September 30, 2026

What AIG and Deutsche Telekom Know About AI Orchestration

When agentic workflows fail, most teams reach for a better model or another agent. See why AIG, Walmart, and Deutsche Telekom focused instead on which agent does what and when a human steps in.

6 min

  • 33% of high-severity AI agent incidents cause cascading failures, where one agent's bad output travels down a chain.

  • AI orchestration is what keeps agentic workflows reliable as the number of agents grows.

  • To scale multi-agent AI, organizations can set a clear handoff order, build a shared orchestration layer, and draw an escalation boundary for each action.

Paul Estes

Editor-in-Chief | AI Advisory | Fractional COO

Rachel Lewis manages thousands of people at BNY, America's oldest bank. She also manages 9 digital employees who never sleep, never take a sick day, and don't even have names.

These digital employees are AI agents, but BNY treats them a lot like human staff. Each one has a login, an email address, a human manager, and a specific role with clear expectations. Some validate payments, find the cause of a stuck transaction in minutes rather than hours, and others repair code.

In fact, Rachel's nine are only a fraction of the roughly 140 digital employees BNY had by early 2026. Meanwhile, more than half of its human staff were building AI agents too. That kind of growth could easily turn into agent sprawl, where agents multiply faster than anyone can track who owns them or what they can reach, and agentic workflows stop working. 

To avoid this, BNY took a different route. In one analyst's view, most banks racing toward agentic AI start with the agents. This bank started with the platform and its people. 

Shortly after ChatGPT launched in late 2022, BNY set up an AI Hub and then built Eliza, a platform that connects AI models to its internal data and compliance controls. BNY trained its people at the same pace it built the technology. In 2025 alone, employees completed 171,000 hours of AI learning through bootcamps and self-paced courses.

Only then did digital employees move into core operations, where an agent touching a regulated process needs least-privilege access, continuous monitoring, and a record of every action. At a bank of BNY's size and importance, those controls aren't optional.

"Every action the digital employee performs is fully auditable… the rationale is logged."

— Michael Demissie, Head of Applied AI, BNY

With those controls in place, AI now handles a real share of BNY's regulated work. It reviews about 70% of restricted party screening for payments, checks whether anyone in a transaction appears on a sanctions list, and resolves those cases more than 30% faster. AI also addresses more than 10% of custody settlement inquiries, where investigations now run more than 80% faster.

BNY chart showing about 70% of restricted party screening for payments reviewed by AI with over 30% faster resolution, and over 10% of custody settlement inquiries addressed by AI with over 80% faster investigation.
Source: BNY — AI now reviews a measurable share of BNY's regulated work and resolves it faste

What sets BNY apart isn't that it has AI agents. It's that the bank built the platform first, trained its people next, and added agents last, each one running inside an AI orchestration layer that gives it an owner, a scope, and an audit trail.

So what is an agentic workflow? It's a process where AI agents carry out tasks, make decisions within set limits, and pass work along. When several specialized agents divide that work, it becomes a multi-agent AI system.

In this article, we'll look at the three ways AI agentic workflows break and how leading enterprises fixed each one.

More Agents Don't Make Agentic Workflows Work Better

BNY put its orchestration layer in place before it scaled its agents. Many enterprises are doing the reverse, adding agents faster than the AI orchestration that coordinates them.

The numbers show just how fast agentic AI is scaling. Among organizations with at least $1 billion in revenue, the share scaling AI agents in at least one function has risen from 27% to 40% since last year. By 2028, the average global Fortune 500 enterprise is expected to run over 150,000 agents, up from fewer than 15 in 2025.

The orchestration layer isn't keeping up. In a survey of 501 US leaders involved in their organizations' agentic AI strategy, only 15% had reached scaled, orchestrated, multi-agent adoption. Just 25% expect agents to coordinate across functions within 2 years. For almost everyone else, that coordination is 4 years away or more, as the chart below shows.

Deloitte bar chart showing when 501 surveyed leaders expect process redesign, cross-functional agent coordination, and blurred functional boundaries, with 25% expecting agents to coordinate across functions within 2 years and 58% within 4.
Source: Deloitte — Most leaders expect agentic AI workflows to coordinate across functions within 4 years rather than 2.

‍

Governance faces the same problem. Only 13% of organizations believe they have the right AI agent governance in place, and some respond by blocking or restricting agents. That can push employees toward shadow AI when sanctioned tools are too restrictive.

"Organizations need to find a balance where they can govern agents and manage sprawl, but also safely empower employees to innovate with these tools."
—Max Goss, Senior Director Analyst, Gartner

That balance gets harder as the number of agents grows. A model upgrade changes what an individual agent can do. AI orchestration has a different job. It decides which agent takes each task, what information passes from one agent to the next, and when a human needs to step in.

Next, we'll look at where those workflows actually break.

Three Ways Agentic Workflows Break Before Any Agent Fails

When a multi-agent AI system underperforms, the reflex is to fix the agents with a better model, sharper prompts, or one more agent. But each fix improves a single agent, while the failure usually starts between them.

The cost of weak AI orchestration is already visible. A recent study shows organizations averaged 54 AI agent incidents last year that needed human correction. Of the high-severity incidents, 33% caused cascading system failures, where one bad output traveled down a chain instead of stopping.

Those failures aren't inevitable. Organizations that build control into their AI systems see 25% fewer incidents and run 16 times more agents than those relying on manual governance.

Here's a closer look at the 3 ways agentic workflows break:

  • Handoffs With No Set Order: When agents are layered onto existing processes, they inherit handoffs that were never designed for them, and the output can look finished even when steps were skipped. The chart below shows why that gap is costly. Value rises with each wave of AI and peaks with multi-agent systems, where agents own every step and every handoff runs between them. With 7,000 to 8,000 subprocesses in a typical business, those gaps add up quickly. The fix is to map the handoffs first, give each agent one job, and let a single agent assemble the result.

‍

 

Source: BCG — Multi-agent systems sit at the far end of the curve where agents coordinate across functions and systems.
  • No Shared Orchestration Layer: Individual teams can build agents that work well, but often no one builds what connects them. Users face a growing list of tools with no clear place to start, and an agent that picks up a request midway has no record of what came before. Just over half of large organizations use AI agents, yet only 18% orchestrate multiple agents across workflows. The fix is AI orchestration, a single entry point that routes each request to the right agent while each team keeps ownership of what it built.
  • No Escalation Boundary: An agent flagging a problem can often act alone, while an agent changing a live system may need a person to sign off first, but that line isn't always drawn action by action. The survey results below show the gap. Of 950 business leaders, 74% are giving agentic AI access to their data and processes, yet just 1 in 5 has tested a plan for when something goes wrong. Only 5% let agents make high-stakes decisions without human review. The fix is to set autonomy per action and track each agent's own accuracy and response time, so drift shows up before an incident does.

‍

Source: Grant Thornton — Most leaders give agentic AI access to their data while few have tested what happens when it fails.

‍How AIG Mapped Its Handoffs Before Building Any Agent

Agents layered onto a process nobody has mapped inherit handoffs they were never designed for. AIG faced that risk at scale. In 2024, Lexington, the excess and surplus lines business of global property and casualty insurer AIG, received roughly 300,000 new-business submissions and bound about 6,700.

Each submission arrived in pieces. A single one could include broker cover letters, statements of value, loss runs, and supplemental applications, all in different formats, and underwriters had to read and reconcile every document by hand. A complex submission could take 3 to 4 weeks to review. More underwriters would have added capacity, but they wouldn't have made any single review faster.

AIG's first step was AIG Assist, a generative AI tool that extracts data from those documents, checks each submission against the business's risk appetite, and ranks it for review. In an early deployment, turnaround fell to under a day, and extraction accuracy rose from about 75% to over 90%.

The gains carried further down the workflow. In Lexington's middle-market property business, underwriters quoted 30% more submissions, time to quote fell 55%, and submissions bound rose by about 40%. CEO Peter Zaffino later said that came without additional headcount.

AIG Assist handled the initial groundwork. The next step was to break that work into a multi-agent AI system, with each agent responsible for a defined part of underwriting. That meant revisiting the process itself. 

AIG built a detailed map of how the work gets done, including each step and what information passes from one step to the next. It then used that map to build the AI orchestration layer that coordinates the agents and their handoffs.

With that structure, each agent is designed to do one clear job. As AIG described on its earnings call, one agent ingests submissions and extracts the data, handing structured information to the next. Another evaluates the risk against underwriting guidelines, a third benchmarks pricing against portfolio targets, and a collaboration agent brings those outputs together for the underwriter.

"These agents will communicate and hand off work to each other to augment our underwriters, just like a well-functioning underwriting team, but operating at machine speed and with inherent consistency."
—Peter Zaffino, Chairman and CEO, AIG

The underwriter stays in the chain, and AIG says it can monitor each agent's activity and intervene in real time.

The multi-agent AI system is still in development, but AIG has already set the order. It mapped the handoffs first, gave each agent one job, and let a single agent assemble the result, so its agents won't inherit handoffs that were never designed for them.

How Walmart Built a Shared Orchestration Layer for Its Agents

Individual teams can build agents that work well, but if no one builds what connects them, users end up with a growing list of tools and no clear place to start. Walmart hit that point as its agentic workflows multiplied.

In July 2025, the largest retailer in the US announced a company-wide framework built around 4 super agents, after finding that its growing set of useful agents was becoming harder to navigate.

The agents themselves weren't the problem. Teams had been building them quickly, and 900,000 associates were already asking 3 million questions a week through the company's conversational AI.

The problem grew as the number of tools increased. Customers, associates, suppliers, and developers all had more places to go and no obvious place to start. 

"Once we saw how quickly teams were adopting these agents and how helpful they were, we realized agents weren't just useful; they were essential. But we also recognized that multiple agents - even if each one is useful - can quickly become overwhelming and confusing."
—Suresh Kumar, EVP, Global CTO and CDO, Walmart

Walmart's answer was to give its agentic workflows a common front door. Sparky serves customers, Marty serves suppliers and sellers, an associate agent serves employees, and a developer agent serves technology teams.

The specialized agents underneath stayed where they were. An agent that handles benefits questions still does that job, while the associate agent acts as the single point of entry to every back-end agent.

That handoff needs more than a directory. Element, Walmart's ML platform, added a stateful architecture that tracks what agents say, do, and aim to achieve, with pipelines that carry context to the right step and protocols that let agents coordinate. WIBEY, the developer super agent, reads what a user wants and orchestrates the work across domains, so a single prompt can fix a pipeline without the user knowing which system handles it.

Interestingly, builders kept their agents. Under WIBEY's federated model, domain teams own what they build while the AI orchestration layer keeps those agents discoverable and interoperable. Teams can now build and share small nano agents in as little as a week.

Across the business, Walmart says AI has cut customer support resolution times by up to 40% and shifted team-lead planning from 90 minutes to 30.

The important change wasn't fixing a broken agent. Walmart built a shared orchestration layer to connect the agents it already had. Users now have one clear place to start, agents no longer pick up requests with no record of what came before, and each team still owns what it built. That is how agentic AI workflows scale without turning into a growing list of tools.

How Deutsche Telekom Drew an Escalation Boundary for Its Agents

An agent flagging a problem on a live network can often act alone, but an agent changing that network may need a person to sign off first. Deutsche Telekom drew that line action by action.

Getting far enough to draw it is the hard part. In telecom, agentic AI has mostly stayed at the proof-of-concept stage, and Deutsche Telekom found that moving a system into production took 4 to 5 times the effort.

Its reason to push ahead was a shortage of hours. A team had been tuning the network by hand for about 1,000 events a year across Germany, including concerts and sports matches where mobile demand can spike quickly, and each site took about an hour.

In November 2025, Deutsche Telekom put its RAN Guardian Agent into live network management, which it says made it the first network operator to do so with a highly developed AI agent. Once running, the system found 40,000 relevant events a year, far more than the team had ever covered by hand. But Deutsche Telekom didn't give one system control over everything that followed.

Instead, the team built a multi-agent AI system and split the work across 3 agents. An event agent finds and verifies events and assigns confidence scores. A monitoring agent evaluates capacity needs. A remediation agent changes network configurations in real time as crowds build.

Deutsche Telekom's own diagram shows that handoff in action. The event agent spots a traffic jam on a Rhine bridge, the monitoring agent checks how nearby antennas are coping, and the remediation agent lists three fixes. Only one of them, adjusting an antenna's angle, is marked for operator approval.

Source: Deutsche Telekom — In RAN Guardian's agentic workflow, some actions run on their own while others wait for operator approval.

That separation gave the team a way to set the boundary for each action. About 75% of actions run without human approval, while the remaining 25% need a person to act on the agent's recommendation. Either way, nothing happens out of sight, since the team can see everything the agents are doing. 

"With an intelligent interaction between our network experts and AI, we are solving specific challenges for the benefit of our customers - for the best network. And we are taking a big step towards autonomous, self-healing networks."
—Abdu Mudesir, Board Member for Product and Technology, Deutsche Telekom

Autonomy also needs a way to spot drift. Deutsche Telekom tracks network measures such as throughput and latency alongside agent measures including response time and accuracy.

Managing a major event now takes around a minute instead of hours, a more than 95% improvement, and RAN Guardian triggered more than 100 remediation actions on its own in its first month. During Germany's Carnival season in February 2026, when street parties draw large crowds, it pre-checked 611 mobile sites across about 130 events, and only 5 needed adjusting. The system is now scaling to the Czech Republic and Croatia.

Deutsche Telekom didn't set one autonomy level for the whole system. It set autonomy per action, kept a person signing off on the rest, and tracked each agent's own accuracy and response time, so drift can show up before an incident does. That escalation boundary is what keeps its agentic workflow safe to run on a live network.

Agentic Workflows Break Between the Agents, Not Inside Them

BNY reached about 140 digital employees by putting the platform, training, and controls in place first. Each agent arrived with an owner, a defined scope, and an audit trail.

The same pattern runs through the agentic workflows at AIG, Walmart, and Deutsche Telekom, each closing a different break. AIG gave its handoffs a set order before its agents arrived. Walmart built a shared orchestration layer to connect the agents its teams already had. Deutsche Telekom drew an escalation boundary action by action. 

In each case, the fix sat between the agents, not inside them. That is key as deployments grow. By 2030, 50% of AI agent deployment failures are expected to come from governance platforms that can't enforce rules at runtime and agents that can't work across systems. 

AI orchestration is what closes that gap. It decides what an agent receives, where its output goes next, and when a human steps in. A better model improves what one agent does. It doesn't build those connections.

So the question for any team weighing its next agent isn't whether the model is good enough. It's whether the agentic workflow around it already defines what that agent receives, what it can finish alone, and who gets involved when it can't.

‍

Cut through the AI hype and join the thousands of business leaders getting practical enterprise insights delivered to their inbox

Welcome to the community! We'll be in touch soon.

Frequent Asked Questions

No items found.