From AI pilot to production: a five-step playbook for putting AI into real work
Almost every business now uses AI somewhere, yet few can point to a result on the books. The companies that get there follow the same pattern: one measurable process, redesigned before any model touches it, a pilot with a gate, review rules written in from day one, and scale only for what already pays.
By Monolith

Picture a regional distributor that gave its team AI assistants last spring. People like them. The sales coordinator drafts emails faster, the office manager summarizes meetings, and a few staff swear by it. Six months later the owner asks a simple question: did anything on the profit and loss statement change? Nobody can say yes.
That company is typical. In McKinsey's 2026 State of AI survey, nearly nine in ten respondents report regular use of AI in at least one business function, and eight in ten say it has improved their own productivity. But the share reporting any contribution to earnings is essentially unchanged from a year ago at 37%, and the high performers who get meaningful value from AI remain just 6% of respondents. McKinsey's summary is blunt: individual gains have yet to translate into broad financial impact. 3
The gap is wider for smaller companies. Fifty-four percent of organizations with at least $1 billion in revenue say they are scaling AI across the enterprise, compared with one-third of smaller ones. That is not because small firms lack tools. It is usually because nobody owns the step from a promising demo to a changed process. This playbook is that step. 3
Why AI pilots stall
Microsoft has been unusually candid about its own rollout. Its lesson is that access and usage do not equal transformation: a tool licensed to over 200,000 people does not change how the work gets done. Its second lesson is sharper. Adding agents to a broken process still leaves a broken process, and speeding up one step just creates a longer queue at the next. 1
Those two sentences explain most stalled pilots. The pilot proves that a model can draft, summarize or classify. It does not prove that the business process around it will absorb the output, that anyone measured the before and after, or that someone is accountable when the model is wrong. The five steps below address each of those gaps in order.
- Pick one measurable processHigh volume, repeated weekly, with a number you already track.GateA baseline is written down: time, volume, errors or cost today.
- Redesign the work firstMap every step, cut waste, then decide where AI fits.GateA new process map shows which steps disappear and which AI takes.
- Run a bounded pilotOne team, one KPI, a fixed number of weeks.GateThe KPI is moving against the baseline, not just usage.
- Design review in from day oneDecide who is accountable, what is logged and what triggers a human.GateThree governance questions answered in writing.
- Scale only what paysExpand to more teams or processes once the numbers hold.GateA value dashboard shows baseline, current reading and target.
Step 1: pick one process you can measure
Start from the business outcome, not the tool. When Microsoft's own sales rollout plateaued, the team stopped pushing adoption harder and started from the business goals, mapping how account managers actually spent their week. IBM recommends a similar exercise: ask each function to estimate the dollar impact of its top three AI opportunities. The point is to rank candidates by value before anyone opens a chatbot. 14
Good first candidates share three traits. They happen often enough that small gains add up, such as quote requests, invoice matching or order status questions. They follow a repeatable pattern. And they already produce a number: turnaround time, cost per order, error rate, or hours per week. If you cannot write down today's number, you will not be able to prove tomorrow's improvement.
Step 2: redesign the work before adding a model
This is the step most companies skip, and it is the clearest difference between winners and everyone else. McKinsey found that nearly three-quarters of high performers fundamentally redesigned workflows because of their AI use, compared with just one-quarter of other respondents. They rebuild the process around what AI makes possible rather than inserting AI into the old one. 3
IBM puts it bluntly: adding AI to an inefficient process automates the dysfunction. Its suggested test is to map the current workflow end to end and ask whether, designing it from scratch with AI available, you would build it the same way. Microsoft's Frontier Playbook calls the same idea de-waste before automating. 24
Microsoft's cloud supply chain team is the strongest worked example. It simplified its processes first, then deployed more than 100 purpose-built agents across planning, sourcing, fulfillment and logistics. Across five monthly planning cycles, average cycle time fell from about 10 business days to less than 2.5, and demand-plan investigations that took five to seven days dropped to a few hours, some under 20 minutes. Microsoft describes the headline as up to 75% in selected workflows, a best case rather than an average. 1
Step 3: run a bounded pilot with one KPI
A bounded pilot has a fixed team, a fixed length and one number that decides whether it continues. Resist measuring success by how many people used the tool. Usage is the first signal, not the result. Microsoft's playbook sets out a measurement sequence: adoption shows within weeks, process indicators within months, performance after that and business outcomes last. 2
- 2 to 4 weeksInput: adoption and usageFirst signalAre the right people using it on the right task?
- 1 to 3 monthsThroughput: process indicatorsLeading indicatorIs the work moving faster or with fewer handoffs?
- 3 to 6 monthsOutput: productivity and performancePerformanceIs each person or team producing more, or better?
- 6 months or moreOutcome: business impactBusiness resultRevenue, margin, cost or customer results that show up in the books.
The playbook adds a useful diagnostic. If usage looks healthy but process metrics have not moved after three to six months, the problem is habit formation, not the tool. Microsoft's own sales example shows why patience matters: sellers who used the tools regularly had 20% higher close rates and 9.4% more revenue per account manager than low-usage sellers. That compares heavy and light users from 2024 data rather than proving cause, but it is the kind of leading indicator worth tracking. 12
Step 4: design human review in from day one
Review and governance belong in the pilot, not bolted on at the end. IBM's rule is simple: before any AI system goes into production, answer three questions in writing. Who is accountable for this system's outputs? How will decisions be traced and audited? What triggers a human review? Governance added after deployment, it warns, is both expensive and fragile. 4
Microsoft's playbook adds the principle that makes review practical: AI autonomy should be tuned to the stakes of each decision, not set globally. Some steps stay human-led, some become agent-operated with human review, and some may eventually run more autonomously. In the supply chain example, agents act only within defined permissions and approval thresholds. 12
| Decision | Stakes | Autonomy level |
|---|---|---|
| Tag and route incoming emails | Low: easy to fix | Agent operates, spot-check weekly |
| Draft a quote from past jobs | Medium: affects margin | Agent drafts, person approves |
| Reconcile invoices to orders | Medium: money moves | Agent matches, person clears exceptions |
| Refunds, credits and contract terms | High: financial and legal | Human-led; AI only prepares the file |
| Anything that changes a password or payment details | Highest | No AI action at all |

Step 5: scale only what already pays
IBM's scaling tool is a value dashboard that maps every active AI initiative to a specific business metric, with a baseline, a current reading and a target. If an initiative cannot be tied to a metric, either it is not valuable enough to continue or the measurement is not in place yet. Its closing test is worth borrowing: how much of your AI portfolio would you bet real money on delivering business value within twelve months? 4
| Initiative | Metric | Baseline | Current | Target | Decision |
|---|---|---|---|---|---|
| Order status replies | Hours per week | 12 | 5 | 4 | Scale to second location |
| Quote drafting | Days to send quote | 3 | 2.5 | 1 | Fix the process, keep piloting |
| Meeting summaries | No business metric | – | – | – | Keep as a personal tool only |
Adoption is a team sport
Microsoft says an early pilot taught it that tools and training alone were not enough; AI adoption is a team sport. It cites its own survey data showing that when managers actively model AI use, reported value from agentic AI rises 17 points and trust rises 30 points. Those are self-reported figures from a company that sells the tools, which is worth keeping in mind for every Microsoft number in this article. The direction still matches what small teams see: people copy the habits their manager demonstrates. 1
Start this week
- List five repetitive processes and write down one number for each today.
- Pick the one with the highest volume and a number you trust.
- Map it on one page, cross out steps that exist only because of old tools, and mark where AI would act.
- Answer IBM's three questions in writing before the pilot starts.
- Set a four-week pilot with one team, one KPI and a review date on the calendar.
- At the review, scale, fix or stop, based on the number rather than on enthusiasm.
Monolith's take: the companies getting results from AI are not using better models; they are running a better process. Pick one measurable workflow, redesign it before adding AI, gate every step on a number, and write the review rules before launch. Scale the pilots that pay and retire the ones that do not.
Why do most AI pilots fail to reach production?
How long should an AI pilot run?
What should be decided before an AI system goes live?
Read for this feature. The numbers match the markers in the text.
- Microsoft: What we've learned from Microsoft's own AI transformationblogs.microsoft.com
- Microsoft: Becoming a Frontier Firm, our Frontier Playbook (2nd edition)aka.ms
- McKinsey: The state of AI in 2026, on the road to ROImckinsey.com
- IBM: 5 practical ways to scale AI that actually deliver business valueibm.com
Want help choosing the first process, redesigning it and setting up the pilot gates? That is exactly the work we do.
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.