
Who owns the last mile of agentic AI?
We are getting very good at teaching AI when to give up.
That may sound like an odd compliment, given that most of the conversation around agentic AI is currently obsessed with how much more autonomous we can make it. How many tasks can the agent complete? How many systems can it access? How much of the customer journey can it handle before a human needs to get involved?
All sensible questions. But in enterprise customer experience, knowing when NOT to continue is becoming just as important.
Genesys is talking about agentic experience orchestration across AI, data and human support. Salesforce’s Agentforce Contact Center now describes AI-to-human handoffs where the human agent receives the transcript and customer history rather than starting blind.
Which leaves us with another question I don’t think the industry has answered particularly well yet.
Once the system decides a human is needed, who owns what happens next?
A routing decision is only the beginning
Imagine a customer contacts their bank because they have spotted a transaction they don’t recognise. The AI agent does a perfectly respectable job. It identifies the customer, gathers information, checks the obvious possibilities and eventually decides the interaction needs to move to a human. So far, so good. The transcript goes across along with account history. Maybe the reason for escalation goes across too. Technically, the handoff has worked.
Then the human appears and performs the higher-assurance identity verification or KYC. That may mean seeing an identity document or asking the customer to show something that simply cannot be communicated properly through a text box. Perhaps there is a fraud concern where visual verification becomes important, or the conversation has reached the sort of grey area where human judgement is required.
Suddenly, having a transcript is useful, but nowhere near sufficient.
The customer may be sent somewhere else, asked to open another application, repeat security questions or start another process from zero. And this is where our beautiful architecture diagram meets an irritated human being who couldn’t care less which component technically did its job.
They just know they have had to start again.
Context is only one part of a good handoff
There is a lot of discussion at the moment about preserving context between AI and human agents, and rightly so. Making someone explain the same problem twice is maddening. If an AI agent has already spent five minutes gathering information, checking the account and eliminating possible solutions, all of that should follow the customer.
But I think we risk defining “context” too narrowly if we reduce it to a transcript and some CRM fields.
Some interactions contain information that exists outside the conversation itself. You may need to see the person. You may need to verify that the person is physically present and who they claim to be. You may need to inspect a document, a damaged product, a medical issue, a signature or some other piece of visual evidence. There may be regulatory requirements around how that interaction is recorded and audited.
And sometimes the human simply needs access to more than words on a screen because the stakes have changed. That is the point where the AI handoff becomes an infrastructure problem.
The orchestration platform knows when. Someone still has to solve how.
I think this distinction will become increasingly important as agentic systems mature.
Platforms such as Genesys are becoming extremely sophisticated at orchestration. They can work across customer data, workflows, AI agents and human teams to determine what should happen next. Genesys itself now describes agentic AI as moving beyond scripted automation towards intelligent orchestration.
That is exactly what enterprises need, but once the orchestration engine decides that an interaction requires a live human, there is another layer underneath that decision.
How does the customer reach them? Does the interaction stay inside the experience they are already using? Does the human receive the context? Can identity be verified? Can they move directly into live video if the situation requires visual interaction? Can the organisation control where that session is hosted, what data is retained and how the interaction is audited?
With VideoEngager, the interaction can remain inside the experience the customer is already using. Browser-based WebRTC requires no plugin or download and can be embedded directly into an enterprise’s web, mobile, kiosk or contact-centre environment. For regulated enterprises, that can mean peer-to-peer WebRTC, customer-controlled storage, private-cloud or on-premises recording, configurable retention, secure vaulting and controls such as WAF, IP allowlisting, mTLS, privacy and consent management.
These sound like implementation details until you are operating in financial services, healthcare, government or another regulated environment. Then they become rather important details indeed. This is the part of the stack we are focused on at VideoEngager
VideoEngager provides the trusted human-engagement layer between AI automation and high-stakes human interaction. It allows an AI conversation to escalate into live video while preserving the surrounding context and adding the identity, visual and compliance capabilities required when text or voice alone no longer does the job.
In other words, the orchestration system decides that the customer needs a person. We make that decision executable.
This is already happening at scale
One reason I feel strongly about this is that video engagement in regulated customer journeys is well past the experimental stage.
BAC Credomatic, one of Central America’s largest financial institutions, handles hundreds of thousands of VideoEngager interactions annually across six countries. Video is used within real financial-services workflows, including identity verification and KYC. Across those journeys, BAC reports an average effectiveness rate of 80%, an average NPS of 79 and 85% first-contact resolution. Hundreds of thousands of interactions a year changes the conversation somewhat. At that scale, you are no longer debating whether customers will use video as part of a digital journey. You are deciding where visual human engagement belongs inside the architecture and what should trigger it.
Agentic AI makes that question considerably more interesting. If AI can handle more of the routine work, then the interactions reaching humans are likely to become disproportionately complex, unusual, sensitive or valuable. Those are exactly the interactions where dropping somebody into a disconnected channel makes the least sense.
The last mile is where trust becomes operational
There is a temptation with AI to talk about trust as though it were primarily a branding problem. Tell customers when they are talking to AI, put the right disclosures in place, explain how their data is being used, publish an AI policy somewhere sensible. All useful.
But trust also has an operational component. When the machine reaches its limit, can I reach a person without beginning the entire journey again? If identity suddenly matters, can you verify me? If I need to show you something, can I? If this interaction later becomes disputed, is there an appropriate record of what happened? And if your AI has already spent five minutes learning what is wrong, does the human who arrives know any of it?
These are fairly basic expectations from a customer’s perspective. They become much harder when you try to deliver them securely across AI systems, contact-centre platforms, CRM infrastructure, compliance requirements and multiple communication channels.
That is why I think the last mile of agentic AI deserves much more attention.
The next generation of enterprise AI will undoubtedly become better at deciding when humans should enter the loop. The companies that get this right will also think very carefully about what happens in the seconds immediately afterwards, because a smart AI agent followed by a clumsy human handoff still produces a clumsy customer experience.
The customer will not grade the individual components.
They will grade the journey.
Experience It Yourself
Experience the power of VideoEngager AI with a free trial or request a demo now!
