The 60-Second Version

Everyone is talking about "AI agents," but the word hides a huge range. On one end sits a chatbot that writes you a draft and waits. On the other sits a system running on its own server, triggered by a message or a schedule, reading your files and sending emails while you sleep. Both get called "AI." They are not remotely the same decision.

The useful question is not which tool is best. It is how much are you willing to let it do without a human in the loop. Every step you give up on oversight buys you speed and leverage, and costs you control, reversibility, and the ability to explain what happened afterwards.

We find it helps to think of four rungs on a ladder, from AI that only suggests to AI that acts on its own. Understanding what each rung gives you, and what it takes, is how you avoid handing an autonomous agent the keys before you have asked whether it should have them.

The Autonomy Ladder AI suggests, you act AI acts, no human 1 · Chat You drive. AI writes, you execute. 2 · Assisted Acts in one surface. You approve each edit. 3 · Agentic Runs multi-step work. You gate risky actions. 4 · Autonomous Runs on its own, across systems, no human by default.
Figure 1: The autonomy ladder. Capability rises as human oversight falls.

How They Compare

Before walking through each rung, here is the high-level picture. We score the four across the dimensions that actually decide enterprise adoption: how much human oversight each needs, its capability, its blast radius (how much damage a mistake can do), how reversible its actions are, and how enterprise-ready it is today.

Rung Human oversight Capability Blast radius Reversibility Enterprise-ready
1 · ChatEvery actionAdvice, draftsMinimalFullHigh
2 · AssistedEvery editIn-context code and textOne file or surfaceHighHigh
3 · AgenticGated approvalsMulti-step tasksA repo or sessionModerateGrowing
4 · AutonomousLittle to noneCross-system actionEverything it can reachLowEmerging
Key Takeaway

There is no single right rung. The right choice depends on whether the action is reversible, how much oversight your team can realistically sustain, and how much a mistake would cost. Most mature setups use different rungs for different tasks, not one for everything.


Rung 1: Chat

You drive, the AI advises

ChatGPT, Claude.ai, Microsoft Copilot Chat, Gemini

This is where almost everyone starts, and for good reason. You ask a question, the model answers. It can draft an email, explain a contract clause, or write a block of code, but it cannot do anything. Nothing changes in your systems unless you copy the output and act on it yourself.

The safety comes from that gap. The human is the executor of every action, which means a wrong answer is a wasted paragraph, not a deleted file or a sent email. For research, drafting, analysis, and thinking out loud, this is often all you need, and it is the right default for anything sensitive where you are not yet sure you can trust the output.

The limit is equally clear: it does not scale past your own hands. Every action still costs your time, because you are the one carrying the AI's output into the real world.

Rung 2: Assisted

The AI acts, but inside one surface

GitHub Copilot, Cursor, Windsurf, in-app AI assistants

Here the AI starts to act, but only within tight walls. An in-editor coding assistant writes directly into your file, an in-app assistant edits the document you are in. It is taking actions, not just suggesting, but those actions are confined to one surface and you approve each change before it sticks.

This is a genuine step up in leverage. The AI does the mechanical work and you stay in the review seat. The blast radius is one file or one document, and because you are watching every diff, a bad suggestion is caught before it lands.

The catch is subtle: the more fluent these tools get, the easier it is to wave changes through without really reading them. The oversight only protects you if you actually use it.

Rung 3: Agentic, with approval

The AI runs multi-step work, you gate the risky parts

Claude Code, OpenAI Codex CLI, Cline, Gemini CLI

This is the rung that changed the conversation. A tool like Claude Code does not just edit one file. Given a goal, it plans, reads across a codebase, writes to multiple files, runs shell commands, executes tests, and works through a task over many steps. It is genuinely agentic: it decides the sequence of actions itself.

Crucially, the good tools keep a human in the loop through approval gates. The agent proposes to run a command or make a change, and you approve or deny before it proceeds. You are no longer reviewing every keystroke, but you are still the gate on anything that touches the real world. The blast radius is real (a session can rewrite a lot of a project) but it is bounded to the workspace you have given it, and the actions are mostly reversible if you are working in version control.

Watch Out

Approval fatigue is the failure mode here. When an agent asks permission forty times in a session, the temptation is to switch on "auto-approve everything" and walk away. That single toggle quietly moves you from Rung 3 to Rung 4 without you deciding to. Treat that setting as an autonomy decision, not a convenience one.

Rung 4: Autonomous

The AI acts on its own, continuously, across systems

OpenClaw, autonomous agent platforms, always-on custom agents

This is the far end of the ladder, and it is a different kind of thing. An autonomous agent does not wait for you to open it. It runs as a persistent service, triggered by a schedule, a webhook, or a message in a chat app, and it acts across many systems at once: reading your email, editing files, calling APIs, committing code, all with memory that persists between runs.

OpenClaw is the clearest recent example. It is an open-source, self-hosted gateway that connects chat apps like WhatsApp, Slack, and Telegram to AI agents, so you can message it and it goes off and does the work, on your own hardware, around the clock. The appeal is obvious: real leverage, work happening while you are away, one assistant wired into everything you use.

So is the risk. By default there is no human in the loop. The blast radius is everything the agent can reach, which is often a lot, and its actions (a sent message, a deleted record, an external API call) are frequently not reversible. Autonomy is a configuration choice on these platforms: you can require approval before high-risk actions and restrict what it can touch. Whether those guardrails are set up well is the entire question for enterprise use.

The Accountability Question

When an autonomous agent sends the wrong email or changes a customer record at 3am, who is accountable? In a regulated Australian business, "the AI did it" is not an answer your board, your auditor, or the regulator will accept. Autonomy does not remove human accountability, it just moves it away from the moment of action, which makes a complete audit trail non-negotiable.


What Most Teams Get Wrong

Across the deployments we see, the same patterns come up again and again.

Pitfall 1: Skipping to the top of the ladder

The autonomous agent is the exciting one, so teams reach for it first. But most business value sits comfortably at Rungs 2 and 3, with far less risk. Climb the ladder deliberately. Only move to full autonomy for a task once you have proven the agent does it reliably under supervision.

Pitfall 2: Confusing "it can act" with "it should act"

Capability is not permission. An agent being technically able to send emails or push code does not mean that action should run unattended. Decide autonomy per action, not per tool. Reading data can be automatic; anything irreversible or externally visible should stay gated until you have real reason to trust it.

Pitfall 3: No audit trail and no undo

The higher you climb, the more you need to reconstruct what the agent did and why, and to reverse it. Teams deploy autonomous agents with neither. Before granting autonomy, insist on logged actions, human-readable reasoning, and a rollback path. If you cannot answer "what did it do and can we undo it," you are not ready for that rung.

Pitfall 4: Ignoring the agent's credentials

An autonomous agent acts using credentials, and teams routinely hand it broad, standing access to "make it work." That is a serious exposure: a compromised or confused agent can do anything its keys allow. Give agents narrowly scoped, revocable permissions, and never more access than the specific task requires.


The Australian Governance Angle

For regulated organisations here, autonomy is not just an engineering decision, it is a compliance one. Two threads matter most.

Operational risk and accountability. Under APRA's CPS 230, regulated entities are expected to manage operational risk across their critical operations, including material technology and the third parties behind it. An autonomous agent acting inside a critical process is squarely in scope. You need to show the controls, the oversight, and a clear line of human accountability, which pushes hard toward gated approvals and thorough logging rather than unattended action.

Privacy exposure. The moment an agent can read across your systems, it can touch personal information, and Privacy Act obligations follow it wherever it goes. An autonomous agent wired into email, files, and customer records concentrates that exposure in one place. Scope what it can access, and keep a record of what it actually accessed.

Reality Check

None of this rules out agentic AI. It rules out unaccountable agentic AI. A gated Rung 3 agent with scoped permissions and full logging can be both compliant and genuinely useful today. An unsupervised Rung 4 agent with broad access is where most of the regulatory and reputational risk lives.


How to Choose

Rather than comparing every tool, start each task with two questions.

Question 1: Is the action reversible?

If the worst case is a wasted draft or a code change you can revert in version control, you can afford more autonomy. If the action sends something to a customer, moves money, or changes a system of record, it should stay gated behind a human until you have strong, tested reason to trust it.

Question 2: Can you sustain the oversight that rung needs?

Every rung demands attention: reviewing every edit, approving every risky command, or auditing an agent's overnight activity. Be honest about what your team will actually keep doing. Autonomy you cannot supervise is not leverage, it is unmonitored risk.

Decision flow: how much autonomy? Pick a task Is the action reversible? No Keep it gated Rung 1 to 3, human approves Yes Can you sustain the oversight? No Yes Consider autonomy Rung 4, scoped and logged
Figure 2: Two questions decide the rung. Reversibility sets the ceiling, sustainable oversight sets the floor.

Where to Start

If you are working out how much autonomy to give AI in your organisation, a simple sequence works well.

Start at the rung that matches the risk, not the hype. For most knowledge work, chat and assisted tools deliver real value with almost no downside. Prove the value there before climbing higher.

Introduce agentic tools with the gates on. Claude Code and its peers are ready for serious work today, used with approval prompts and version control. This is where most teams should be investing right now: high leverage, bounded blast radius, reversible actions.

Treat full autonomy as a graduation, not a starting point. Move a task to an autonomous agent only once it has earned trust under supervision, with scoped credentials, logging, and a rollback path in place. Autonomy is something an agent earns per task, not a switch you flip once.

Bottom Line

The question is never just "chat or agent." It is "how much should this specific action run without me." Match the autonomy to the reversibility of the action and the oversight you can sustain, and the tooling decision mostly makes itself.

OZ Data Solutions

Founder and Principal Consultant. PhD in computer science, ten years building data and AI systems across eight sectors. More about the practice