ai:production
Scaling AI from pilot to production.
By Loan Laux · September 28, 2026 · 9 min read
Scaling AI from pilot to production means taking a workflow that worked in a demo and making it survive real volume, real permissions, and real accountability: a named owner, defined inputs, a failure path, monitoring, and a budget line. Most pilots never make that jump, and rarely because the model underperformed. They stall because the pilot was built by an enthusiast on sample data, nobody agreed up front what number it had to move, and the unglamorous wiring (access, error handling, ownership) was never anyone's job.
This post walks the jump in order: why pilots stall, what production actually means, the gates a workflow must pass before rollout, and the maintenance that keeps it alive afterward. It's the same sequence we run when a client brings us a promising pilot, whether we embed an engineer to do the build or carry the workflow long-term through managed AI operations.
Why AI pilots stall.
The numbers on this are blunt. MIT's Project NANDA report, The GenAI Divide: State of AI in Business 2025, found that about 95% of enterprise generative AI pilots delivered no measurable P&L impact. Gartner predicted in mid-2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs, and unclear business value. Whatever the exact figure at your organization, the shape is familiar: an impressive demo, a quarter of enthusiasm, then quiet abandonment.
The autopsy is usually the same regardless of industry. The pilot ran on hand-picked sample data, so the first week of real inputs broke it. It was built and babysat by one enthusiast, so it stopped the week they got busy. Nobody agreed before the pilot what number it had to move, so nobody could say whether it worked. And the boring parts (data access, permissions, what happens when the output is wrong) were deferred to 'after we validate the concept', which is exactly when momentum runs out.
Notice what's absent from that list: model quality. The gap between pilot and production is organizational, and no model release fixes an ownership problem.
Pilot versus production.
It helps to be precise about the two states, because most stalled pilots live in an undeclared middle: too load-bearing to kill, too fragile to trust. The differences that matter:
| Dimension | Pilot | Production |
|---|---|---|
| Runs on | Hand-picked sample data | Whatever the real workflow produces, including the ugly cases |
| Runs when | When the champion has time | Every time the trigger fires, no human required to remember |
| When it fails | Someone shrugs and does it manually | A defined fallback path, and someone is notified |
| Owned by | Whoever built it, informally | A named owner with time budgeted for upkeep |
| Measured by | Demos and enthusiasm | A metric agreed before launch, reviewed on a cadence |
| Paid for | A personal or team card | A budget line someone defends at planning |
Scaling AI from pilot to production, step by step.
01:
Pick the one pilot worth scaling
Score your pilots the way an AI workflow audit scores candidates: frequency times time saved, discounted by error cost and data sensitivity, divided by the lift to productionize (the full scoring method is public). Scale the winner. Park the rest without guilt; a parked pilot costs nothing, a half-scaled one costs attention forever.
02:
Agree the number before you build
Pick the single metric the workflow must move (hours per ticket, turnaround days, error rate) and measure two weeks of the current process as a baseline. This is the step teams skip, and it's why so many pilots end in an argument about whether they worked.
03:
Re-scope from demo to workflow
A demo answers 'can the model do this?'. A workflow spec answers what triggers the run, what goes in, what must never go in, what checks the output passes, and where it lands. Write that on one page before touching infrastructure; every later step falls out of it.
04:
Wire the real data and permissions
Connect the actual systems: the ticket queue, the document store, the CRM, with service accounts and access scoped to what the workflow needs. This step is the most common graveyard. If IT approval takes six weeks, start this step first, not last.
05:
Build the failure path
Decide what happens when the model output is wrong, the API is down, or an input arrives malformed: who gets notified, what the human fallback is, and what gets logged. A workflow without a failure path doesn't fail gracefully; it fails silently, and silent failures are what get AI projects banned by leadership.
06:
Write the SOP and train the people in the loop
The reviewers, the fallback humans, and the owner all need to know their part without asking the builder. One page per workflow: trigger, inputs, the AI step, checks, owner, version. We covered the exact format in AI SOPs for agencies, and it applies unchanged here.
07:
Run both processes in parallel, then cut over
Two to four weeks of the AI workflow running alongside the manual one, comparing outputs against your baseline metric. Cut over when the numbers say so, not when the demo feels good, and only then retire the manual path.
What keeps a production workflow alive.
Launch is the midpoint. The ground moves under a production AI workflow in a way it doesn't under normal software: providers update and retire models on their schedule, not yours; input patterns drift as the business changes; volume grows past what the pilot ever saw; and the people in the review loop turn over. The maintenance load is real, and it's the half of the work that pilots never surface:
- Output monitoring. Sample real outputs weekly against the checks in the SOP. Quality regressions rarely announce themselves; they show up as a slow slide reviewers stop noticing.
- Model change management. When a provider updates or deprecates a model, someone re-runs the evaluation set before the swap, not after users complain.
- Cost tracking. Per-run cost times real volume, reviewed monthly. Workflows that were cheap at pilot volume have a way of becoming line items worth renegotiating.
- The metric review. The number from step two, on a monthly cadence, with authority to tune or kill. A workflow nobody measures anymore is a workflow nobody can defend at budget time.
- A named owner with budgeted time. Every item above loses to deadline pressure the week it belongs to nobody. This ongoing carry is precisely what managed AI operations exists to take off your plate.
Where scaling goes wrong.
The failure modes past the pilot stage are fewer, but more is riding on them:
- Scaling ten pilots at once. Every workflow you productionize needs the same wiring, review, and maintenance. Organizations that scale one deep build the muscle; organizations that scale ten wide build a backlog of half-owned automations.
- Keeping the champion as the architecture. If the person who built the pilot is still the only one who can fix it in production, you haven't scaled, you've just renamed the pilot. The SOP and the parallel-run period exist to break exactly this dependency.
- Treating review as a launch phase. Teams often plan generous human review 'until we trust it', then quietly drop it. Review levels should be set by error cost in the SOP, permanently, not by how novel the workflow feels.
- Building everything internally on principle. The MIT report found externally built tools reached production roughly twice as often as internal builds. Build what touches your differentiated data and process; buy or adapt the rest.
Do it yourself, or bring us in.
Everything above is doable internally if you have an engineer who can own the wiring for a few weeks and, harder, an owner who keeps the workflow honest after launch. The pattern we see when it stalls anyway: the build competes with roadmap work and loses, or launch succeeds and maintenance belongs to nobody by Q2.
We work both halves of the problem. Our forward-deployed engineers embed with your team to take the pilot through the steps above and into production, and managed AI operations carries the monitoring, model changes, and metric reviews afterward, so the workflow still earns its budget line a year later. If you're not yet sure which pilot deserves this treatment, start with an audit and let the scoring decide.
Frequently asked questions.
Why do most AI pilots fail to reach production?
Rarely because of the model. MIT's State of AI in Business 2025 report found about 95% of enterprise generative AI pilots showed no measurable P&L impact, and the causes it and Gartner point to are organizational: no agreed success metric, pilots built on sample data by a single enthusiast, poor data quality, and no plan for permissions, failure handling, or ownership. Those are all fixable before the pilot starts, which is the cheapest place to fix them.
How long does it take to move an AI pilot into production?
For one workflow with a cooperative IT department, typically six to twelve weeks: a couple of weeks to spec and baseline, a few for data access and the build, and two to four for the parallel run against the manual process. Data access and permissions are the usual long pole, which is why we start that step first. Anything promised faster is usually skipping the parallel run.
Should we scale one pilot or several at once?
One, almost always. The first production workflow teaches you where your data access is slow, what your review culture actually tolerates, and what maintenance really costs. Those lessons make the second and third workflows dramatically cheaper. Running the sequence in parallel across many pilots means learning every lesson the expensive way, several times, simultaneously.
Do we need to hire AI engineers to do this?
You need engineering time, not necessarily a hire. The production wiring (data connections, failure paths, evaluation) is weeks of a capable engineer's attention per workflow, plus a fraction of someone's time for upkeep. Some organizations pull that from an internal platform team; others embed one of ours precisely because the roadmap always outbids internal AI plumbing. What doesn't work is expecting the pilot's original champion to carry production alone from the side of their desk.
How do we know if the production workflow is actually working?
You compare against the baseline you measured before building, on the one metric you agreed up front, reviewed monthly by the workflow's owner. If you skipped the baseline, measure two weeks of the current state now, before cutover; without it you're arguing from anecdotes at budget time.