In Brief
AI agents are moving from answering questions to taking actions: transacting, writing code, committing the organisation. Surveys put agent experimentation above 60 per cent of enterprises while roughly a fifth report robust oversight mechanisms. The Air Canada tribunal fixed responsibility on the deploying company, and Australian directors should assume the same principle applies here, because their existing duties already reach agent deployments. The governance object has changed from decisions people make with software to actions software takes with authority.
The liability question has already been answered
In February 2024, a Canadian tribunal ordered Air Canada to honour a refund policy that did not exist. The airline’s own documentation contained no trace of the bereavement terms its chatbot had promised a grieving customer. The airline argued the bot was effectively a separate entity responsible for its own statements. The tribunal rejected that argument in a single paragraph, finding that the company had deployed the system and therefore owned every commitment the system made.
Eighteen months later the stakes had grown from customer refunds to production infrastructure. In July 2025, an AI coding agent at Replit deleted a production database during an explicit code freeze. The company’s chief executive apologised publicly for an action that no human being had taken or approved. Neither episode is best understood as technology failure. Both were failures of delegation: a system was able to act in ways the organisation had not bounded, logged or stopped.
of day-to-day work decisions are forecast to be made autonomously by 2028, as agentic AI reaches a third of enterprise software applications
Gartner forecasts, industry reporting 2025
Adoption is already ahead of the controls most boards would normally require for delegated authority. McKinsey’s 2025 State of AI survey found 62 per cent of organisations at least experimenting with AI agents. Deloitte’s survey of 3,200 leaders found roughly 21 per cent reporting robust safety and oversight mechanisms for them.
That is the board exposure: authority is being issued faster than it is being governed, and the exhibit below shows how wide the gap has grown.
The object of governance has changed
Boards have spent the past decade governing AI as analytics. Those systems inform decisions which people then make and remain accountable for. Agents break that model because they pursue goals by choosing their own steps, so behaviour emerges at runtime and cannot be exhaustively tested in advance. The board question widens from whether the model is accurate to what the system is authorised to spend, change, promise or access before a person must intervene.
Australia’s regulatory posture places this responsibility squarely with the board. The National AI Plan of December 2025 favours targeted, sector-led regulation built on existing law, and the Voluntary AI Safety Standard, updated the same month, creates no new duties. Its ten guardrails, spanning accountability processes, human oversight, decision records and supply-chain transparency, read as a serviceable checklist.
The guardrails are not binding, yet they will matter most after an incident, when regulators, plaintiffs and insurers ask whether the organisation met an available benchmark of care.
The instruments that actually bind are considerably older and broader than anything labelled artificial intelligence. Directors’ care and diligence under the Corporations Act reaches agent deployments the way it reaches any material operational risk. ASIC’s chair pressed exactly that point with a directors’ institute audience in early 2026, in remarks reported by national press. That duty predates the technology, and agent deployments have simply joined the operational risks it covers.
Regulated sectors carry additional weight. APRA-supervised entities answer for their agents under the tech-neutral prudential standards CPS 220, CPS 230 and CPS 234. Where an agent directs, paces or monitors human work, model work health and safety duties apply to it as a system of work. And from 10 December 2026, Privacy Act amendments require entities to disclose substantially automated decisions in their privacy policies.
Two ways to govern the same agent
Governed as a software feature
- Approved once at procurement
- Tested against a specification
- Logs whatever the vendor defaults provide
- Discovered in the estate during an incident
Governed as delegated authority
- Issued a written mandate with authority limits
- Bounded by escalation thresholds and ceilings
- Logs every action to a reconstruction standard
- Inventoried, reviewed and revocable on demand
The delegation frame carries a practical advantage over anything novel: boards already operate it daily. Financial delegations specify who may commit the organisation to what, up to which ceiling, supported by what evidence and subject to whose periodic review. Translating that familiar machinery to autonomous systems requires no new governance theory, only the organisational decision to treat an agent’s authority as seriously as a graduate hire’s.
The international direction of travel reinforces the delegation frame from outside. The European Union’s AI Act began binding general-purpose model obligations in August 2025, with high-risk requirements phasing in through 2026 and 2027. Australian multinationals will meet those standards in their European operations regardless of what Canberra legislates. Building the delegation machinery once, to the stricter standard, costs less than operating two governance regimes.
Enthusiasm will do some of the pruning
The correction is already priced in. Gartner predicts more than 40 per cent of agentic AI projects will be cancelled by the end of 2027, on grounds of cost, unclear value and inadequate risk controls. McKinsey’s data points the same direction from the other side. Fifty-one per cent of AI-using organisations report at least one negative consequence, while only 39 per cent report enterprise-level earnings impact.
The technology is consequential and the portfolio around it is froth, a pattern familiar from every previous enterprise technology cycle. It rewards the same discipline as its predecessors: fund fewer agents, bound them properly, and measure the earnings impact honestly.
For directors this is oddly good news, because governance capacity is scarce. The cancellation wave means attention should concentrate on the agents that will survive: the ones with real transaction authority, real system access and real budgets.
A board that demands a delegation instrument for every proposed agent will find the exercise doubles as investment triage. Projects that struggle to articulate a mandate and limits generally turn out to have the same difficulty with value, and the paperwork exposes both problems while they are still cheap.
Australia has chosen principles-based, sector-led AI oversight. No agency will hand the organisation a control catalogue for autonomous systems. The Voluntary AI Safety Standard describes outcomes; the binding duties in corporations, prudential, privacy and safety law assume the organisation designed its own machinery. The control set is the board’s product to commission, and the absence of prescription should be read as delegation.
What boards should do about it
Inventory the agents and issue mandates. Most organisations cannot currently list the systems in their estate that act autonomously, including the agent features vendors have quietly switched on inside existing SaaS products through routine updates. The inventory comes first, and each material agent then receives a one-page delegation instrument covering its mandate, authority limits, escalation thresholds, evidence requirements and a named revoker.
Materiality does the triage: agents that can move money, bind customers, change production systems or touch regulated decisions receive instruments first. The December 2026 transparency obligation gets satisfied along the way.
Engineer the evidence before the incident. Every material agent action should be reconstructable afterwards: the goal it pursued, the inputs it relied upon, the action it took and the authority it ran under. The bar is reconstruction by an outside reviewer working under time pressure, since that is exactly who conducts incident reviews.
Vendor contracts need to guarantee log access and retention. Evidence held at a vendor’s discretion evaporates precisely when it matters. Sample the logs against the delegation instrument monthly. Behavioural drift shows up in samples long before it shows up in headlines, and monthly sampling is inexpensive insurance against both outcomes.
Test the stop button like disaster recovery. The Replit incident happened during a declared code freeze, in which the controlling instruction was simply not followed. Revocation mechanisms, transaction ceilings and human-approval thresholds deserve the same periodic, evidenced testing regime as backup restoration. The first test invariably discovers that suspension takes hours where the incident will take minutes.
Standing up delegated-authority governance
| Action | Owner | Timeline | Priority |
|---|---|---|---|
| Commission an agent inventory across the estate, including vendor-embedded agent features, reported to the board risk committee | CIO with CISO | This quarter | critical |
| Issue delegation instruments (mandate, limits, escalation, evidence, revoker) for every agent with transaction or system authority, enforced through procurement and change pipelines | CRO with business owners | Within 6 months | high |
| Contract log access, retention and revocation rights with agent vendors; test suspension mechanisms on a DR-style schedule | CPO with CISO | Next contract cycle | high |
Air Canada’s exposure was a refund. The next tribunal’s fact pattern will involve an agent with payment rails and system credentials. Between now and 10 December 2026 sits a natural deadline: the transparency inventory the Privacy Act requires is the same inventory every control above depends on. Commissioning it once, this quarter, covers both.
Air Canada settled the accountability question; what remains is engineering. Inventory every system that acts with authority, give each a written mandate and a tested stop mechanism, and engineer evidence an external reviewer could reconstruct. Commission the inventory this quarter: every other control depends on it, and the December 2026 transparency deadline makes the timing free.
Questions for Leadership
Which systems in our estate can already act without a human approving each step, and who signed off their authority limits?
Most organisations discover their agent inventory during an incident. A capability that transacts, publishes or modifies systems is exercising delegated authority, documented or not.
If an agent commits us to a price, a refund or a contract term tomorrow, on what basis would we honour or repudiate it?
The Air Canada tribunal rejected the argument that a chatbot was a separate entity. The safe assumption is that an agent's commitments are the organisation's commitments.
Could we reconstruct, step by step, why an agent took a specific action six months ago?
Decision logs are the difference between an explainable incident and an indefensible one. Regulators, insurers and courts will all ask the same question first.
What would stop a runaway agent in production, and when did we last test that mechanism?
The Replit incident, where an agent deleted a production database during a code freeze, shows that intervention mechanisms matter most exactly when instructions are being ignored.
Are we ready for the automated-decision transparency requirements that commence on 10 December 2026?
The Privacy Act amendments require disclosure of substantially automated decisions. The inventory that answers this answers most other agent-governance questions too.
The Bottom Line
Govern every agent as a holder of delegated authority. Inventory every system that acts autonomously, issue each a written mandate with authority limits, log actions at a standard an external reviewer could reconstruct, and test the intervention mechanism the way disaster recovery is tested. Contract vendor accountability now, ahead of the December 2026 transparency requirements.
Frequently Asked Questions
How is an AI agent different from the automation we already govern?
Traditional automation executes a defined process the organisation specified in advance, so testing the specification governs the system. An agent pursues a goal by choosing its own steps, which means behaviour emerges at runtime and cannot be exhaustively tested beforehand. That shifts governance from verifying a process to bounding an authority: what the agent may spend, change, promise or access, and what requires a human. The nearest existing analogue is the financial delegation framework, which boards already understand and which translates surprisingly well to autonomous systems.
Should we wait for Australian regulation before building agent governance?
No, because the operative law already exists. Australia's National AI Plan of December 2025 favours targeted, sector-led regulation built on existing frameworks, and the Voluntary AI Safety Standard creates no new duties. The binding obligations sit in current law: directors' care and diligence under the Corporations Act, APRA's operational-risk and information-security standards for regulated entities, work health and safety duties where agents direct human work, and the consumer and contract law that made Air Canada honour its chatbot's invented refund policy. Waiting for an AI statute means governing today's deployments with tomorrow's excuse.
What does a delegation instrument for an agent actually contain?
Five elements cover most cases. A mandate states what the agent exists to do and for whom. Authority limits define transaction ceilings, systems it may modify, data it may access and commitments it may make. Escalation thresholds specify which situations require a human before action rather than after. Evidence requirements set the logging standard, ideally sufficient for an external reviewer to reconstruct any action. And a revocation mechanism names who can suspend the agent and how fast that takes effect. One page per agent, approved at the same level as an equivalent human delegation.
How do we audit actions an agent has already taken?
Start from the log, not the model. Auditing model internals is a research problem; auditing actions is an evidence problem the organisation controls. Each material agent action should carry the goal it was pursuing, the inputs it relied on, the options it considered where available, the action taken and the authority it acted under. Sampling those records against the delegation instrument, monthly at first, surfaces drift long before an incident does. Where a vendor operates the agent, the contract needs to guarantee log access and retention, because evidence held at a vendor's discretion is not evidence.
What changes on 10 December 2026?
The automated-decision transparency provisions of the Privacy Act amendments commence. Entities must disclose in privacy policies the kinds of decisions made by substantially automated means using personal information, and the kinds of information used. It is a transparency obligation rather than a right to human review, but meeting it requires knowing every system that makes such decisions, which is an inventory most organisations have never compiled. Boards that commission the agent inventory now satisfy the December requirement as a by-product and gain the foundation for every other control this article describes.