← All posts

ai:roi

Measuring AI ROI at your agency.

Measuring AI ROI at an agency means comparing a measured baseline of how a workflow ran before AI against the same metrics after, workflow by workflow: hours per occurrence, turnaround, and error rates, minus what the automation costs to build, run, and review. Most agencies skip the baseline, which is why most AI ROI claims are a feeling with a percentage sign attached. The method below replaces the feeling with a spreadsheet row per workflow.

The measurement discipline is the back half of our AI workflow audit method: the same scoring that decides what to automate also defines the baseline and the one metric each workflow has to move, so ROI is designed in before the build rather than reconstructed after it.

Why most AI ROI numbers are vibes.

MIT's Project NANDA report, The GenAI Divide: State of AI in Business 2025, found that about 95% of enterprise generative AI pilots delivered no measurable P&L impact. The load-bearing word is measurable. Plenty of those pilots probably did something useful; nobody could prove it, so at budget time they counted as nothing.

Agencies produce the same gap in smaller rooms. The evidence offered for 'AI is working here' is usually seat counts on a tool subscription (adoption, not return), a survey where the team feels faster (feelings, not hours), or a time-saved estimate multiplied by a rate card (a guess multiplied by a bigger number). None of these survive a skeptical finance question.

The fix is not more sophisticated accounting. It's measurement hygiene: measure the workflow before, measure it after, subtract the costs, and write the result down where a partner can challenge it.

The model: baseline, delta, cost.

Per workflow, ROI is the value of what changed minus what the change cost. Three quantities, all measurable:

  • The baseline. Two representative weeks of the current process, logged before anything is built: hours per occurrence summed across everyone who touches it, elapsed turnaround, and how often the output comes back with errors or revisions. Boring to collect, and the single highest-value artifact in this whole exercise.
  • The delta. The same metrics once the workflow is live and past its shakedown period, from the same sources. Not 'the team says proposals feel faster': hours logged on proposals, days from discovery call to sent proposal, revision rounds per deliverable.
  • The cost. Everything on the other side of the ledger: subscriptions and API usage at real volume, the build and tuning time, and the ongoing human share, meaning review passes, exception handling, and maintenance. The review tail is the line agencies most often forget, and it's frequently the largest.

Measuring AI ROI, step by step.

  1. 01:

    Pick the workflows, not the company

    ROI only computes at the workflow level: proposals, reporting, content production, QA. If you haven't chosen which workflows to automate yet, run the scoring from our audit method first; the same logic that picks winners tells you where measurement is worth the effort.

  2. 02:

    Baseline before you build

    Two representative weeks of the current process, logged as it happens. Hours per occurrence across every role that touches it, elapsed turnaround, and the error or revision count. Reconstructed baselines inherit everyone's optimism; live ones don't.

  3. 03:

    Agree the one number per workflow

    Before launch, the workflow's owner and whoever controls the budget agree on the single metric it must move and by roughly how much. This is cheap insurance against the most common ending, which is an argument about whether the thing worked.

  4. 04:

    Price the saved hours honestly

    An hour saved is only worth money if it turns into billable work, absorbed growth, or an actual payroll change. We price redeployed hours at realized rates and unredeployed ones near zero, and the numbers stay defensible.

  5. 05:

    Count the whole cost side

    Subscriptions and usage at real volume, the build and tuning time at the builder's real cost, and the permanent review share. If a senior person now spends five hours a week checking outputs, those hours belong in every month's denominator, not just the first one.

  6. 06:

    Review monthly, with authority to kill

    The workflow's owner reviews the metric against baseline on a monthly cadence, with real authority to tune or shut it down. A workflow nobody measures anymore is a workflow nobody can defend, and a negative-ROI workflow nobody can kill is a subscription with ceremony.

What to measure, workflow by workflow.

The right metrics are mostly obvious once you commit to the per-workflow frame. For the workflows agencies automate most:

WorkflowBaseline to captureDelta to watch
ProposalsHours per proposal; days from call to sendTurnaround days, plus win rate over a full quarter
Client reportingPM hours per reporting cycleHours per cycle; corrections issued after sending
Content productionDrafting hours per piece; revision roundsHours per piece; client revision requests
Meeting follow-upTime from call to tickets and recapRe-entry time; action items that get dropped
First-pass QARework hours; defects caught per stageEscaped defects; rework hours per project

Quality and error deltas, not just hours.

Hours are the easy half of the measurement, and alone they mislead. If the automated version is faster but worse, the cost shows up later as rework, client corrections, and churn risk, none of which appear in a time log. Four quality signals are trackable with almost no tooling:

  • Revision rounds. How many passes a deliverable takes before sign-off, internal and client-side. The first number to move when output quality drops.
  • Escaped errors. Corrections issued after a deliverable reached the client. Rare, expensive, and the one metric where a small negative delta can erase a large positive one.
  • Review burden. Minutes of senior review per output. If automation moved the work from drafting to fixing, the hours didn't disappear, they changed desks.
  • Spot-check scores. A weekly sample of outputs scored against the workflow's checklist. Quality regressions rarely announce themselves; a slow slide only shows up if someone is sampling.

The returns a spreadsheet misses.

Some real returns resist the per-workflow math, and pretending otherwise produces exactly the inflated decks this post argues against. Faster turnaround wins deals, and over a quarter that shows up in win rate, a number you can watch even if you can't cleanly attribute it. Absorbed growth is similar: taking on more clients without hiring is visible in revenue per head over quarters, not in a weekly time log.

The margin story matters most. As we argued in will AI replace agencies, AI is repricing agency work either way, and under hourly billing your efficiency gains become your client's discount. Measured ROI is what makes the alternative work: you can only price fixed-scope work confidently when you know what delivery actually costs you. Report these second-order effects as observed trends next to the per-workflow numbers, and resist the urge to assign them invented dollar values.

Where measurement goes wrong.

  • Measuring after the fact. 'How long did proposals take before?' asked in month three gets answers shaped by whoever is proudest of the project. Everything else in this method tolerates improvisation; the baseline doesn't.
  • Company-level math. Blending every tool and workflow into one AI ROI figure produces a number nobody can act on. Strong workflows subsidize weak ones and the weak ones never get killed.
  • Rack-rate hours. Valuing every saved hour at the billable rate assumes every freed hour becomes billable work. It doesn't, and finance knows it doesn't, which is how the whole analysis loses credibility.
  • Ignoring the review tail. Launch-month costs get counted; the permanent review and maintenance share doesn't. The parallel-run method in scaling AI from pilot to production exists partly to surface this cost before cutover.
  • Survivorship reporting. Only the workflows that worked make the slide. Killed experiments cost real money and belong in the denominator; counting them honestly is also what makes the successes believable.

Do it yourself, or bring us in.

Nothing above needs software you don't already own: a time log, a shared sheet, and a monthly calendar slot per workflow. The hard part is discipline, specifically the two-week baseline before anyone gets to build the fun thing, and a review cadence that survives busy months. That's also where internal efforts quietly fail: the baseline gets skipped in the excitement, and by the time someone asks for the ROI, the 'before' no longer exists.

Our AI workflow audit builds the measurement in from the start: the audit's scoring sets the baseline and the target metric for each workflow before anything is automated, and the roadmap ships with owners and checkpoint dates attached. For workflows already in production, managed AI operations carries the monthly metric reviews so the numbers keep existing after the novelty wears off. Either way, the deliverable is the same: an AI line in the budget that can defend itself.

Talk to us about it

Frequently asked questions.

How do you calculate AI ROI for an agency?

Per workflow: the value of the measured deltas (hours saved at an honest rate, turnaround compression, error reduction) minus the full cost (subscriptions and usage, build time, and ongoing review and maintenance hours), compared against a baseline measured before the automation went in. Company-wide AI ROI is just the sum of these rows.

We automated months ago and never measured a baseline. Is it too late?

You can't recover the real before, but you can stop the bleeding: measure the current state now as the baseline for everything that changes from here, and for one or two workflows, run a small manual sample alongside the automated path for two weeks to approximate the old cost. It's cruder than measuring first, and it's still far better than reconstructing hours from memory.

How long before AI ROI shows up?

Time-based deltas are readable about a month after a workflow stabilizes, once the parallel-run period is done. Quality deltas need a quarter, because escaped errors and revision patterns are lumpy. Win rate and revenue per head need two or three quarters of trend. Any AI ROI claim made in week two of a rollout is a projection, whatever the slide calls it.

Is counting hours saved enough to prove ROI?

No, for two reasons. Saved hours only have value if they're redeployed into billable work, absorbed growth, or changed staffing, so the pricing has to be honest. And hours say nothing about quality: if revision rounds or client-visible errors rose, the rework can consume the savings. Hours plus a quality delta plus full costs is the minimum credible package.

What ROI should an agency expect from AI?

Distrust benchmark percentages; they're mostly marketing, and the honest answer is that returns are concentrated. In the agencies we work with, a small number of high-frequency workflows (proposals, reporting, meeting re-entry) carry most of the measurable return, while long-tail experiments roughly break even. That concentration is the argument for scoring workflows before automating rather than spreading tools everywhere and hoping.

Keep reading.

:

How to run an AI workflow audit at your agency.

:

Will AI replace agencies? No, but it will reprice them.

:

Scaling AI from pilot to production.