InsightsPublished 9 min read

From AI pilot to production: a five-step playbook for putting AI into real work

Almost every business now uses AI somewhere, yet few can point to a result on the books. The companies that get there follow the same pattern: one measurable process, redesigned before any model touches it, a pilot with a gate, review rules written in from day one, and scale only for what already pays.

By Monolith

An operations manager and two colleagues at a whiteboard process map, one crossing out a step, with a warehouse visible through the window
An operations team crossing a step off its process map before adding AI to it.

Picture a regional distributor that gave its team AI assistants last spring. People like them. The sales coordinator drafts emails faster, the office manager summarizes meetings, and a few staff swear by it. Six months later the owner asks a simple question: did anything on the profit and loss statement change? Nobody can say yes.

That company is typical. In McKinsey's 2026 State of AI survey, nearly nine in ten respondents report regular use of AI in at least one business function, and eight in ten say it has improved their own productivity. But the share reporting any contribution to earnings is essentially unchanged from a year ago at 37%, and the high performers who get meaningful value from AI remain just 6% of respondents. McKinsey's summary is blunt: individual gains have yet to translate into broad financial impact. 3

The gap is wider for smaller companies. Fifty-four percent of organizations with at least $1 billion in revenue say they are scaling AI across the enterprise, compared with one-third of smaller ones. That is not because small firms lack tools. It is usually because nobody owns the step from a promising demo to a changed process. This playbook is that step. 3

Why AI pilots stall

Microsoft has been unusually candid about its own rollout. Its lesson is that access and usage do not equal transformation: a tool licensed to over 200,000 people does not change how the work gets done. Its second lesson is sharper. Adding agents to a broken process still leaves a broken process, and speeding up one step just creates a longer queue at the next. 1

Those two sentences explain most stalled pilots. The pilot proves that a model can draft, summarize or classify. It does not prove that the business process around it will absorb the output, that anyone measured the before and after, or that someone is accountable when the model is wrong. The five steps below address each of those gaps in order.

The pilot-to-production playbook
  1. Pick one measurable processHigh volume, repeated weekly, with a number you already track.GateA baseline is written down: time, volume, errors or cost today.
  2. Redesign the work firstMap every step, cut waste, then decide where AI fits.GateA new process map shows which steps disappear and which AI takes.
  3. Run a bounded pilotOne team, one KPI, a fixed number of weeks.GateThe KPI is moving against the baseline, not just usage.
  4. Design review in from day oneDecide who is accountable, what is logged and what triggers a human.GateThree governance questions answered in writing.
  5. Scale only what paysExpand to more teams or processes once the numbers hold.GateA value dashboard shows baseline, current reading and target.
Monolith's five-step sequence, drawing on Microsoft's published lessons, its Frontier Playbook, McKinsey's high-performer findings and IBM's governance guidance. Each gate must be passed before the next step. [1][2][3][4]

Step 1: pick one process you can measure

Start from the business outcome, not the tool. When Microsoft's own sales rollout plateaued, the team stopped pushing adoption harder and started from the business goals, mapping how account managers actually spent their week. IBM recommends a similar exercise: ask each function to estimate the dollar impact of its top three AI opportunities. The point is to rank candidates by value before anyone opens a chatbot. 14

Good first candidates share three traits. They happen often enough that small gains add up, such as quote requests, invoice matching or order status questions. They follow a repeatable pattern. And they already produce a number: turnaround time, cost per order, error rate, or hours per week. If you cannot write down today's number, you will not be able to prove tomorrow's improvement.

Step 2: redesign the work before adding a model

This is the step most companies skip, and it is the clearest difference between winners and everyone else. McKinsey found that nearly three-quarters of high performers fundamentally redesigned workflows because of their AI use, compared with just one-quarter of other respondents. They rebuild the process around what AI makes possible rather than inserting AI into the old one. 3

IBM puts it bluntly: adding AI to an inefficient process automates the dysfunction. Its suggested test is to map the current workflow end to end and ask whether, designing it from scratch with AI available, you would build it the same way. Microsoft's Frontier Playbook calls the same idea de-waste before automating. 24

Microsoft's cloud supply chain team is the strongest worked example. It simplified its processes first, then deployed more than 100 purpose-built agents across planning, sourcing, fulfillment and logistics. Across five monthly planning cycles, average cycle time fell from about 10 business days to less than 2.5, and demand-plan investigations that took five to seven days dropped to a few hours, some under 20 minutes. Microsoft describes the headline as up to 75% in selected workflows, a best case rather than an average. 1

Step 3: run a bounded pilot with one KPI

A bounded pilot has a fixed team, a fixed length and one number that decides whether it continues. Resist measuring success by how many people used the tool. Usage is the first signal, not the result. Microsoft's playbook sets out a measurement sequence: adoption shows within weeks, process indicators within months, performance after that and business outcomes last. 2

When each signal shows up
  1. 2 to 4 weeksInput: adoption and usageFirst signalAre the right people using it on the right task?
  2. 1 to 3 monthsThroughput: process indicatorsLeading indicatorIs the work moving faster or with fewer handoffs?
  3. 3 to 6 monthsOutput: productivity and performancePerformanceIs each person or team producing more, or better?
  4. 6 months or moreOutcome: business impactBusiness resultRevenue, margin, cost or customer results that show up in the books.
Measurement stages and indicative timelines from Microsoft's Frontier Playbook, second edition, September 2026. Your own timing will vary with the process. [2]

The playbook adds a useful diagnostic. If usage looks healthy but process metrics have not moved after three to six months, the problem is habit formation, not the tool. Microsoft's own sales example shows why patience matters: sellers who used the tools regularly had 20% higher close rates and 9.4% more revenue per account manager than low-usage sellers. That compares heavy and light users from 2024 data rather than proving cause, but it is the kind of leading indicator worth tracking. 12

Step 4: design human review in from day one

Review and governance belong in the pilot, not bolted on at the end. IBM's rule is simple: before any AI system goes into production, answer three questions in writing. Who is accountable for this system's outputs? How will decisions be traced and audited? What triggers a human review? Governance added after deployment, it warns, is both expensive and fragile. 4

Microsoft's playbook adds the principle that makes review practical: AI autonomy should be tuned to the stakes of each decision, not set globally. Some steps stay human-led, some become agent-operated with human review, and some may eventually run more autonomously. In the supply chain example, agents act only within defined permissions and approval thresholds. 12

DecisionStakesAutonomy level
Tag and route incoming emailsLow: easy to fixAgent operates, spot-check weekly
Draft a quote from past jobsMedium: affects marginAgent drafts, person approves
Reconcile invoices to ordersMedium: money movesAgent matches, person clears exceptions
Refunds, credits and contract termsHigh: financial and legalHuman-led; AI only prepares the file
Anything that changes a password or payment detailsHighestNo AI action at all
Monolith's example of tuning autonomy to stakes for a small business, following the three levels in Microsoft's Frontier Playbook. 2
A team lead and an analyst at a conference table reviewing results on a laptop while the lead points to a printed bar chart
A team lead and an analyst checking a pilot's results against the baseline.

Step 5: scale only what already pays

IBM's scaling tool is a value dashboard that maps every active AI initiative to a specific business metric, with a baseline, a current reading and a target. If an initiative cannot be tied to a metric, either it is not valuable enough to continue or the measurement is not in place yet. Its closing test is worth borrowing: how much of your AI portfolio would you bet real money on delivering business value within twelve months? 4

InitiativeMetricBaselineCurrentTargetDecision
Order status repliesHours per week1254Scale to second location
Quote draftingDays to send quote32.51Fix the process, keep piloting
Meeting summariesNo business metric–––Keep as a personal tool only
An illustrative value dashboard for a small distributor. The figures are invented to show the format, not results from a Monolith client. Format from IBM's guidance. 4

Adoption is a team sport

Microsoft says an early pilot taught it that tools and training alone were not enough; AI adoption is a team sport. It cites its own survey data showing that when managers actively model AI use, reported value from agentic AI rises 17 points and trust rises 30 points. Those are self-reported figures from a company that sells the tools, which is worth keeping in mind for every Microsoft number in this article. The direction still matches what small teams see: people copy the habits their manager demonstrates. 1

Start this week

  • List five repetitive processes and write down one number for each today.
  • Pick the one with the highest volume and a number you trust.
  • Map it on one page, cross out steps that exist only because of old tools, and mark where AI would act.
  • Answer IBM's three questions in writing before the pilot starts.
  • Set a four-week pilot with one team, one KPI and a review date on the calendar.
  • At the review, scale, fix or stop, based on the number rather than on enthusiasm.
The short version

Monolith's take: the companies getting results from AI are not using better models; they are running a better process. Pick one measurable workflow, redesign it before adding AI, gate every step on a number, and write the review rules before launch. Scale the pilots that pay and retire the ones that do not.

Questions, answered
Why do most AI pilots fail to reach production?
Usually because AI is added to an unchanged process and nobody measures the business result. McKinsey finds that high performers are about three times as likely as others to redesign workflows around AI.
How long should an AI pilot run?
Long enough to see process indicators move, often one to three months. Microsoft's playbook expects adoption signals in two to four weeks and business outcomes after six months or more.
What should be decided before an AI system goes live?
IBM recommends answering three questions in writing: who is accountable for the outputs, how decisions will be traced and audited, and what triggers a human review.
Sources

Read for this feature. The numbers match the markers in the text.

  1. Microsoft: What we've learned from Microsoft's own AI transformationblogs.microsoft.com
  2. Microsoft: Becoming a Frontier Firm, our Frontier Playbook (2nd edition)aka.ms
  3. McKinsey: The state of AI in 2026, on the road to ROImckinsey.com
  4. IBM: 5 practical ways to scale AI that actually deliver business valueibm.com

Want help choosing the first process, redesigning it and setting up the pilot gates? That is exactly the work we do.

Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.