← All posts

ai:commerce

AI for ecommerce agencies: what running a commerce shop taught us.

AI for ecommerce agencies works differently than it does for a generalist shop, because commerce delivery runs on structured, repetitive material: catalogs, order flows, theme code, support tickets, promo calendars. That structure is exactly what current models handle well, so the automatable share of a commerce agency's week is higher than most. The catch is that the same work sits closer to a customer's wallet, so a wrong output isn't an awkward draft, it's a wrong price or a broken checkout.

We're not writing this from the outside. Before out:grow became an AI practice, it was a headless commerce agency, founded in 2018 and acquired by Trellis in 2021. This post is what that history taught us about where AI genuinely fits commerce delivery, and the same sequencing logic we now apply in AI workflow audits.

Where this experience comes from.

out:grow started in 2018 as a Dubai-based headless commerce agency built on Reaction Commerce, the API-first platform that became Mailchimp Open Commerce after Mailchimp acquired it. We shipped custom GraphQL APIs and Next.js storefronts for enterprise clients across Europe and Southeast Asia, work most agencies avoided because it needed real engineering rather than a plugin marketplace. In 2021, Trellis, a Boston ecommerce agency, acquired out:grow to stand up its headless practice, and the founder ran that division, then engineering across the combined team after Zaelab acquired Trellis.

One honest footnote from that era: the platform we specialized in was eventually discontinued. Betting on hard technology sometimes means the technology loses. What survived the bet was the capability, because clients were never really buying Reaction Commerce, they were buying a team that could make an unusual stack work in production. Keep that in mind for everything below; it's the same dynamic AI is now running on commerce agencies.

Why AI for ecommerce agencies is a different problem.

A brand agency's raw material is loose: briefs, decks, opinions. A commerce agency's raw material is structured to a degree most industries would envy. Products have schemas. Orders have states. Tickets have categories. Themes have templates. Promotions run on a calendar that repeats every year. Language models are at their best when the task is 'transform structured input into structured output following patterns you've seen before', which describes an enormous share of commerce delivery.

The flip side is error cost. A generalist agency's worst-case AI failure is an embarrassing draft that a human catches in review. A commerce agency's worst case is a hallucinated spec on a live product page, a discount applied to the wrong collection, or copy that contradicts what's actually in the box. The work is more automatable and less forgiving at the same time, which is why sequencing matters more here than anywhere else we've applied AI.

The commerce workflows worth automating first.

These are the candidates that consistently score highest when we run commerce agencies through an audit. All of them share the same shape: high frequency, structured input, and a human checkpoint that's cheap relative to the hours saved.

:

Catalog and product data work

Drafting and normalizing descriptions, attributes, and metadata across hundreds or thousands of SKUs from supplier feeds and spec sheets. The single biggest hour-sink in most commerce retainers, and the most mechanical.

:

Migration data mapping

Translating categories, attributes, and content between platform schemas during replatforms. Tedious, structured, and historically done by developers who hate it.

:

Theme and boilerplate code

First passes at sections, templates, and integration glue. The judgment lives in review and architecture, not in typing out another product-card variant.

:

Storefront QA sweeps

Checking variants, currencies, breakpoints, and promo states against a checklist before launch. Consistent machine sweeps first, human judgment on what they surface.

:

Support triage

Classifying and drafting responses to where-is-my-order and returns tickets, which dominate commerce inboxes and follow a handful of patterns.

:

Reporting and merchandising prep

Assembling the weekly trading report from data that already exists in the platform, ad accounts, and email tool. Pure reassembly.

What we would keep human.

The line we draw is not 'AI drafts, humans review everything equally'. Some work gets automated to the edge of publish; some never leaves human hands regardless of how good the models get:

  • Anything that touches price. Discounts, promo logic, shipping rules. A wrong number here has an immediate dollar cost, and models are confidently wrong about numbers in exactly the way that hurts.
  • Unattended writes to production. Drafting into a staging state is fine. An agent publishing directly to a live storefront is a checkout incident waiting for a busy Friday.
  • Merchandising and UX judgment. Which products deserve the homepage, how a collection should be structured for this brand's buyer. The model can summarize the data; the call is taste plus accountability.
  • Hero copy and brand voice. Catalog long-tail descriptions are pattern work. The twelve sentences that define how a brand sounds are not, and clients can tell the difference faster than agencies think.

The platforms are coming for the easy layer.

Shopify now ships AI into the merchant admin itself: Shopify Magic generates product descriptions and marketing copy from a few product details, and Sidekick acts as an assistant that answers questions about the store and performs admin tasks. That matters strategically. If a chunk of your billable service is 'we write your product descriptions', you're now competing with a feature the platform includes in the subscription your client already pays for.

This is the commerce-specific version of the repricing argument we made in will AI replace agencies: the production layer collapses in price first. What the platform's AI can't do is see across systems. It doesn't know what's in the client's PIM, why the ERP feed disagrees with the storefront, how the email flows relate to the promo calendar, or what the support backlog says about a product page. Cross-system workflows, data quality, and integration judgment are where a commerce agency's AI work stays defensible, for the same reason our headless work was defensible: it's the part that needs someone who can make an unusual stack work in production.

How to sequence it: a commerce-specific scoring pass.

The method is the same one in our AI workflow audit guide: shadow real delivery, list recurring tasks, score them, sequence two per quarter with owners and checkpoints. For commerce agencies we weight two axes differently.

First, error tolerance dominates. Score anything customer-visible or price-touching brutally low, whatever the hours saved, and let internal-facing work (mapping, drafts into staging, QA sweeps, report prep) rise to the top of the first quarter. Second, frequency in commerce follows the catalog and the calendar, not the org chart: a 3,000-SKU client with weekly drops generates more automatable volume than five brochure-site retainers combined. Rank by where the volume actually is, and be honest about the platform question while you're at it. If Shopify is likely to ship a workflow as a native feature within a year, don't build a service line on it; build on the workflows that span systems the platform can't reach.

Do it yourself, or bring us in.

Everything above is runnable internally: pick your two biggest commerce retainers, shadow a delivery week, score the candidates with the error-tolerance weighting, and automate the top two internal-facing workflows first. The usual failure mode is not ability, it's that the people who understand the delivery work are the same people buried in it.

We run this as an AI workflow audit with commerce agencies, and the history matters here: we've sat on your side of the replatform, the promo-calendar crunch, and the catalog deadline, not just the AI side. The deliverable is a sequenced roadmap your existing team executes.

Talk to us about it

Frequently asked questions.

Which ecommerce workflows should an agency automate with AI first?

Internal-facing, high-volume, structured work: catalog data enrichment into a staging state, migration data mapping, first-pass theme code, QA sweeps, and support triage. They save real hours while a cheap human checkpoint catches errors before anything reaches a live storefront. Price logic and unattended publishing come last, if ever.

Should we build on Shopify's native AI or our own tooling?

Use the native features (Shopify Magic, Sidekick) where they're good enough, and don't bill for what they give away. Build your own tooling only for the cross-system workflows the platform can't see: PIM-to-storefront pipelines, ERP reconciliation, support-to-product-page feedback loops. That's also where the defensible service revenue is.

Is AI-generated product copy safe to ship for clients?

For catalog long-tail, yes, with two conditions: every claim is grounded in the actual spec sheet rather than the model's guess, and a human reviews before publish. We treat hallucinated product attributes as the top risk, because a wrong material or dimension on a live page becomes a return, a complaint, or worse. Hero and brand-defining copy we keep human.

Does this differ for headless builds versus template Shopify work?

The workflows are the same; where the hours come back differs. Headless teams get more from code-side automation (API glue, storefront components, integration tests) because they write more custom code per project. Template-driven teams get more from the content and operations side: catalog work, QA sweeps, reporting. Score your own delivery rather than copying another shop's list.

What about customer data in AI tools?

Order and customer records are personal data, so keep them out of general-purpose AI tools unless you've done the vendor diligence and have the client's data-processing terms covering it. Most of the high-value workflows above (catalog, code, QA, reporting) don't need customer-level data at all, which is a good reason to start there.

Keep reading.

:

How to run an AI workflow audit at your agency.

:

Will AI replace agencies? No, but it will reprice them.

:

AI proposal automation for agencies.