August 4, 2026

J&J, Guardian Life, Webster: Scaling AI by Doing Less

J&J ran 900 AI pilots, then cut most of them. Webster Bank stops 9 in 10 ideas before a line of code. Scaling AI is about what you refuse to build, not what you can.

8 min read

  • Only 25% of organizations have moved 40% or more of their AI pilots into production.

  • The ones that do reach production build less. J&J kept up to 15% of 900 use cases that produced 80% of the value. Webster Bank shipped a dozen of 150.

  • The build is not where scaling AI is decided. Teams that reach production pick one workflow, set the number it has to move, and agree in advance on what would end it.

Staff writer

From AI to FinOps, our team's collective brainpower fuels this blog.

A medical student at Stanford Health Care asked a question in 2023 that sounded almost too simple. Why can't you just ask the medical record a question the way you'd ask ChatGPT? That question eventually led to ChatEHR, an AI assistant that helps clinicians search and summarize electronic health records.

Say a patient arrives in the emergency room with chest pain. As Dr. Jonathan Chen, a Stanford hospital physician, puts it, the chest pain isn’t the point. The whole story is prior surgeries, medications, side effects, everything that led to that moment. 

That history can span years of records, and for transferred patients it arrives as hundreds of pages with minutes to read them. Physicians spend up to 60% of their time on administrative work instead of patient care, and chart review is one of the biggest reasons.

The data was already there. Finding what was crucial inside it was the problem.

So Stanford kept the tool small. ChatEHR did chart review, history summaries, and transfer decisions - nothing else - and rather than build a separate application, the team put it inside the Epic system clinicians already used. 

The first AI pilot ran with 33 physicians, nurses, physician assistants, and nurse practitioners. Emergency physicians spent 40% less time reviewing charts during critical patient handoffs.

"It really does free up time, but it also creates much more accurate summaries."

—Michael Pfeffer, Chief Information and Digital Officer, Stanford Health Care

The AI tool worked. But Stanford still didn’t scale it.

Not until it cleared the organization's responsible AI guidelines, written to keep automation that added risk without benefit out of patient care. ChatEHR reached production only in 2025. Stanford started scaling AI only after the system earned it.

Which raises the harder question: what is scaling for AI when the majority of enterprise projects never leave the pilot stage? Not a bigger model or a wider rollout. 

Scaling AI is the move from isolated pilots to production AI that changes a business number. The organizations that make it to production start with a single measurable workflow, define success before development begins, and know what would stop them from scaling.

In this article, we'll break down why most AI pilots stall and how three leading organizations made scaling AI work from pilot to production.

Scaling AI Is a Decision Problem, Not a Model Problem

Most enterprise AI teams have more pilots running than Stanford did, better funding than Stanford had, and nothing in production. What separates them is a set of scaling AI decisions made before deployment.

In fact, Deloitte's 2026 State of AI in the Enterprise report found that access was never the constraint, but the gap from AI pilot to production is. Only 25% of more than 3,000 executives surveyed said their organizations had moved AI pilots into production at a rate of 40% or higher, even though 54% expect to within three to six months. The chart below shows how few have cleared that bar.

A bar chart showing that only 25% of organizations have moved 40% or more of their AI experiments into production.
Source: Deloitte — Most organizations struggle to move from AI pilot to production.

The returns reflect it. Nearly 74% of AI's economic value is concentrated in 20% of organizations. AI is widely available, but the business value isn’t.

So what makes the difference? Organizations seeing the biggest returns are twice as likely to pursue workflow redesign before AI deployment, rather than adding tools to existing processes, and 1.7 times as likely to have a responsible AI framework in place.

The reality is that many AI project management efforts still operate without that discipline. Nearly a third of IT leaders say unclear ROI metrics are holding AI back, and fewer than half have formal success metrics.

"We have more discipline around business cases than most companies. By the time we put something into production, it's been through a series of proof of concepts, there's been a deep dive on financials, and we are able to move quickly and demonstrate value."

—Sean McCormack, Chief Information Officer, First Student

A separate survey of 950 business leaders points the same way. The organizations pulling ahead are scaling fewer AI pilots with clearer exit criteria, not more.

Three Decisions Separating Scaling AI From Stalled Pilots

When an AI pilot stalls, teams usually blame the model, the data, or the budget. The numbers point somewhere else.

In a survey of 782 infrastructure and operations leaders, only 28% of AI use cases fully succeeded and met ROI expectations, while 20% failed outright. The biggest obstacle wasn't the technology. It was projects that were too ambitious or poorly scoped. 

The organizations seeing the best returns treat AI project management as a business discipline, connecting AI to real workflows and starting with a problem they can actually solve. 

Scaling AI comes down to three decisions:

  • Narrow to One Measurable Workflow: Scaling AI fails when organizations try to solve too many problems at once. In BCG's Build for the Future study of 1,000 companies, leaders prioritized an average of 3.5 AI initiatives against 6.1 for everyone else, and anticipated 2.1 times greater ROI. The chart below shows that gap. Fewer projects mean more time to redesign workflows, measure results, and prove value.

A chart comparing AI use cases prioritized by leading companies against others, showing 3.5 versus 6.1, alongside 2.1 times greater anticipated ROI. 
Source: BCG — Companies scaling AI successfully prioritize fewer use cases and anticipate higher returns.

  • Set a Baseline Before the Build: AI pilots often begin without a definition of success, so teams measure whether the model generates good responses instead of whether it improves the business. Nearly 8 in 10 organizations now use generative AI in at least one function, yet 60% still report no enterprise-wide impact on EBIT. Organizations that reach production AI define value before development begins and track it against business outcomes.
  • Define a Stop Condition Up Front: Some AI pilots fail because no one decides when to stop. In interviews with experienced data scientists and engineers on why AI projects fail, 84% pointed to leadership decisions rather than technical limitations, most often priorities that shifted every few weeks, leaving projects abandoned before they could prove value. Teams scaling AI decide up front what would stop a project rather than letting it drift.

What these three decisions have in common is timing. They all happen before the build, not after the pilot stalls. 

Next, we'll walk through how three leading enterprises applied each of these decisions, and what scaling AI actually required of them.

How Johnson & Johnson Cut 900 AI Pilots Down to the Ones That Paid Off

When AI efforts spread across hundreds of experiments, it's difficult to know which ones deserve to grow. The fix isn't a more capable model. Scaling enterprise AI starts by narrowing the work to the workflows that can prove business value and stopping the rest.

Johnson & Johnson, the world's largest pharmaceutical company by revenue, with $88.8 billion in 2024 sales, spent about 3 years encouraging teams across its Innovative Medicine and MedTech businesses to experiment with generative AI. 

Nearly 900 AI use cases emerged across sales, drug development, supply chain, clinical trials, and internal operations under a central governing board. The goal wasn't to scale everything. It was to learn where AI could make a meaningful difference.

The results made the next decision easier. Johnson & Johnson found that just 10% to 15% of AI use cases delivered 80% of the value. It redirected resources to those projects, cut the rest, and shifted ownership from the central board to the business teams responsible for the work.

"We've moved from a thousand flowers to a prioritized focus on gen AI."

—Jim Swanson, Executive Vice President and Chief Information Officer, Johnson & Johnson

The company evaluated each surviving project on ease of implementation, usefulness across the business, and the value it could deliver. The projects that met those standards moved forward.

Today, those projects are bounded AI workflows rather than open experiments. For instance, a sales copilot within the CRM helps representatives prepare for conversations with healthcare professionals using medically approved information. 

Meanwhile, AI predicts supply chain disruptions before they become shortages. Similarly, internal assistants answer employee questions about benefits and company policies. Also, in clinical trials, sites recommended by J&J's AI models enrolled patients at up to 2.6 times the rate of others.

The first phase enabled Johnson & Johnson to identify where AI could be most effective. The next step was deciding which ideas deserved to become part of the business and to make scaling AI possible.

Why Guardian Life Set a Number Before It Built Anything

Scaling AI often stalls because teams can’t agree on what success looks like at the pilot stage. Without a baseline, a project can appear successful simply because the model works, while the business impact is unclear.

Guardian Life, a workplace benefits and planning provider with 2024 revenue of $14.5 billion, was expanding its use of data and AI to improve customer experience, reduce operating costs, and increase employee productivity. To turn those broad goals into decisions, Guardian's Data and AI team built AI project management around a value-tracking framework. 

Initiatives started as hypotheses developed with business leaders, moved through testing and business case development, and only advanced toward scaling once the case held. Each solution also cleared risk, legal, compliance, and cybersecurity review before approval.

That framework pointed Guardian at a workflow it could already measure. In its employee benefits business, preparing a request for proposal and quote took about a week. That became the baseline for the pilot.

The AI solution cut the process to 24 hours, and Guardian is moving it from pilot to scale in 2026.

Guardian didn’t start with the technology and search for a result later. It started with a business outcome, measured the change, and scaled only after the numbers showed the value. The baseline turned scaling AI into a decision, not just an experiment.

The Threshold That Stopped Most of Webster Bank's AI Ideas

Most AI ideas shouldn't reach production. Without a clear threshold for saying no, they consume budget and engineering time without delivering business value. 

The fix isn't a better model. It's deciding in advance what would stop a project before anyone commits resources to scaling AI.

Webster Bank, a commercial bank with approximately $86 billion in total consolidated assets, had no shortage of AI ideas. As it expanded generative AI across document processing, unstructured data analysis, and peer credit reviews, teams across the business kept proposing new ones. The challenge wasn't finding opportunities. It was deciding which ones were worth pursuing.

Webster built the stop condition into its governance process. Every AI use case is reviewed by an AI Governance Committee that reports to the Technology Committee of the Board, bringing together leaders from technology, data, cybersecurity, privacy, compliance, legal, and audit. Before development begins, each project defines measurable KPIs, and the committee evaluates its scalability against the bank's risk profile before it can advance.

"We identify KPIs early and test them throughout the process. If you can't measure, you don't know if something is working or if it should even scale."

—Vikram Nafde, Executive Vice President and Chief Information Officer, Webster Bank

The same discipline extends to technology decisions. Webster defines its governance framework before selecting AI tools, letting its AI footprint and associated risks guide what it buys.

The process is intentionally selective. Of roughly 150 AI use cases collected, only about a dozen have reached production. The rest were filtered out before they consumed additional time, budget, and engineering effort.

Webster treated scaling AI as an investment decision, one every proposal had to justify. Most didn't. Only the projects that cleared those review points became production AI.

Scaling AI Is Decided Before the Build, Not After the Pilot

ChatEHR went from a medical student's question to a tool clinicians use every day because the work was narrowed to one job, measured against a clear baseline, and held to a standard before it reached patients. 

Johnson & Johnson, Guardian Life, and Webster Bank reached production the same way. None of them got there by building a smarter model than the pilots that stalled.

The evidence backs that approach. Redesigning workflows around AI has one of the strongest effects on bottom-line impact of any factor tested, and the roughly 6% of companies that qualify as AI high performers are nearly three times as likely as everyone else to have done it.

So the question for any leadership team still presenting stalled AI pilots isn't whether the model is good enough. It's whether the project was ever scoped to one measurable workflow, given a baseline it had to move, and held to a threshold that would stop it.

Cut through the AI hype and join the thousands of business leaders getting practical enterprise insights delivered to their inbox

Welcome to the community! We'll be in touch soon.

Frequent Asked Questions

How do you move AI from pilot to production?

+

You make three decisions before the build: narrow the work to one measurable workflow, set a baseline the AI has to move, and define what would stop it. Skipping the last one is costly. In interviews with experienced data scientists and engineers, 84% pointed to leadership decisions rather than technical limitations as the reason AI projects fail.

What is production AI?

+

Production AI has cleared the pilot stage and runs as a dependable part of the business, held to a defined outcome and monitored over time. Reaching it is rare: in a survey of 782 infrastructure and operations leaders, only 28% of AI use cases fully succeeded and met ROI expectations, with failures traced to overly ambitious or poorly scoped projects rather than weak models.

Can AI do project management?

+

AI can track progress, surface risks, and measure outcomes, but the decisions that reach production are human: scoping to one measurable workflow, setting a baseline, and defining when to stop. Discipline is the differentiator, not the tool. Nearly a third of IT leaders say unclear ROI metrics are holding AI back.

What is scaling AI?

+

Scaling AI is the move from isolated pilots to production AI that changes a business number, not just a demo of what a model can do. The gap is wide: only 25% of more than 3,000 executives surveyed said their organizations had moved AI pilots into production at a rate of 40% or higher.