Contact us
Train switches

The conversation around AI has moved fast. We kicked off with chat interfaces, then copilots, then agents, and now everything is agentic, a platform, or apparently “the OS” for something.

Most of the noise stops right there. The assumption seems to be that once the agents are good enough, organisations will just chuck them into the business and productivity will magically appear. For enterprise legal ops, compliance teams, financial services and law firms though, that is only half the story (and not even the interesting half).

The real question is not whether an AI agent can do the work. We know they can. Properly instructed, with the right context, they can do genuinely impressive things. The real question is whether that work can be governed to the standard required in regulated industries without the whole thing turning into an unmanageable compliance nightmare.

The enterprise AI gap

AI tools have broadly split into two camps. On one side you have the productivity multipliers: the assistant-style, lawyer-facing platforms like Harvey, Legora, Claude CoWorker and others. These sit alongside the lawyer and help them move faster, draft better, research quicker, analyse more and generally multiply output. The human is still very much in the driving seat. The work happens inside that environment, guided by the operator. That is a fantastic productivity play.

On the other side you have the transactional layer: platforms that take inputs, apply domain expertise, enforce process and push work through to outputs as part of scaled legal and operational workflows. A lot of this is still best handled deterministically. It has to be. The human-in-the-loop is often providing oversight, exception handling, judgement or accountability rather than doing every step manually.

Autologyx has always played mainly in that second camp. We build platforms for complex, high-volume legal and operational processes where the stakes are real and “just ask the AI” is not an operating model.

But the line between those two worlds is blurring fast. The same human operators who used to work inside one productivity tool are now increasingly using those tools to interact with and oversee the transactional layer. Workers and machines are starting to blend, and that is where the real governance problem starts to bite.

The issue is not simply whether an agent can produce a useful answer. The issue is what happens when that answer becomes an action inside a live business process. Take an intake process as a simple example. An agent might read the initial instruction, check the client and matter data, pull supporting documents from the DMS, spot that key information is missing, update the workflow record, draft a response and route an exception to the right person. None of those steps is especially exotic in isolation, and some of them are better handled deterministically than by a model. The point is that, once the agent is changing records, triggering workflow steps or deciding what needs escalation, it is no longer just producing a useful answer in a chat window. It is participating in the operating model.

That is a different category of risk, because work in regulated organisations has rules. Who can see what, who can do what, under whose authority, against which policy, in which jurisdiction, with what approval, with what audit trail, and with what ability to reconstruct the whole thing later when a client, regulator, insurer, GC, risk team or understandably irritated partner asks what happened.

This is why governance cannot just mean an AI policy, a committee, or someone saying “guardrails” on a webinar. In this context, governance means control logic embedded in the work itself. It means permission checks that understand the user, matter, client, jurisdiction and process stage. It means data boundaries that are enforced before the agent sees something it should not. It means approval thresholds, escalation routes, audit events, retention rules and policy decisions being part of the workflow rather than bolted on afterwards in a risk review.

Put another way, the control has to sit in the path of the work. Not somewhere adjacent to it. That is the enterprise AI gap: not intelligence, but control.

From AI tools to AI work

The next proper leap in enterprise AI will not come from shinier chat windows or slightly better demos. It will come from embedding agents into governed business processes so they can participate in real work without breaking the operating model around them.

This lines up closely with what Nikki Shaver wrote recently in her excellent piece on the real battleground in legal AI. Nikki was focused on deep legal research: how lawyers can rely on AI-generated analysis only when it is grounded in authoritative, high-quality legal intelligence that can be referenced, cited and stood behind with confidence. That is critical.

At Autologyx, we are coming at the same problem from the other side. Not the research layer, but the work layer: the transactional, operational, high-volume bit of legal and compliance work where agents are starting to replace or augment people in processes that already have structure, controls, approvals and consequences.

The question there is slightly different. How do you hand meaningful work to agents while maintaining the same level of control, consistency and verifiability you would demand from a human team?

In both cases, the model itself is increasingly table stakes. The real differentiator is the harness around it: the context, workflow, data model, permissions, policy layer, audit and operational controls. Without that, you have clever outputs floating around outside the system, waiting for someone to copy, paste, reconcile, check and explain them. That is not transformation. That is moving the mess around.

The goal is not autonomous AI in the abstract. The goal is controlled delegation, which is a very different thing. Enterprise organisations are not looking for agents that can do whatever they like because the demo looked cool. They want agents that can operate inside clear boundaries, under proper authority, with full visibility of what happened and why.

That is especially true in regulated industries where confidentiality, privilege, client duties, operational resilience and risk controls are not optional extras. They are the price of entry.

Why governance matters more as agents get better

Some corners of the market still behave as if governance is the boring bit that slows everything down. That is backwards. Robust governance is what lets you adopt this stuff at scale without ending up in the papers for the wrong reasons.

The smarter and more capable the agent, the more governance you need. An agent that answers a question badly is annoying. An agent that updates the wrong client record, exposes the wrong document, acts outside delegated authority, relies on the wrong data or moves a regulated process into the wrong state is something else entirely.

The organisation cannot just record that “AI was involved” and call that an audit trail. That will not survive contact with a client, regulator, insurer, risk team or partner who wants to know what actually happened. They will need to know which agent acted, which model or models were involved, which user or workflow delegated authority to it, what data it accessed, what system state it changed, what policy was applied, whether an exception path was followed and whether a human approved the outcome. Not as a nice forensic exercise six months later, but as part of the way the work was controlled at the time.

When everything happened inside one tool with a human user at the keyboard, that was manageable. The agentic world makes it much messier. You have chains of agents, models, tools and systems interacting across platforms and parties, and the more those boundaries blur, the harder it becomes to prove authority, context and accountability after the fact. “The agent did it” is not going to be a satisfactory answer when something goes wrong.

MCP and the future of enterprise AI

MCP is getting attention for good reason. It gives agents a more standardised way to connect to enterprise systems, tools and data sources, which matters if agents are going to do anything more useful than sit in a chat window producing commentary.

But exposing a tool is not the same thing as governing its use. MCP does not, by itself, decide whether that agent should be allowed to use that tool under that user’s authority, against that matter, in that jurisdiction, at that stage of the process, with that approval threshold and that audit requirement. That decision still has to come from the operating layer around the work.

That is where authentication, permissions, workflow state, policy enforcement, exception handling, auditability and human intervention actually matter. Not as abstract governance concepts, but as the boring enterprise machinery that determines whether this stuff can be deployed safely.

As organisations move from sandbox experiments into actual deployment, these controls will matter more than the model of the week. Not because models do not matter. Of course they do. But because in regulated environments, capability without control is just another risk surface.

The conversation we should be having

The excitement around agents is fair enough. The potential is real. But we should spend a bit less time asking “what can agents do?” and a bit more time asking “how do we govern the work they actually perform?”

That is where enterprise adoption really happens. Not at the edge of the business, with clever tools producing outputs that sit outside the operating model, but inside the actual workflow, records, permission model, audit trail and systems that already carry the risk.

That is where AI becomes more than a productivity layer. It becomes part of how work gets done. And in legal, compliance, financial services and other regulated markets, that only works if the work is governed.

Agents matter, of course they do. But agents are not the story. Governed AI work is.