Service

AI Orchestration

The useful AI work in commerce is unglamorous: generating the first draft of five thousand product descriptions, categorising support tickets, flagging anomalies in stock data. We build those workflows and keep a person in the approval path.

Diagnosis first

Where AI pays in a commerce operation

Not in a chat widget bolted to the homepage. In the repetitive internal work that currently consumes your team's week.

  1. Thousands of products need copy

    A strong case. Generated first drafts from structured attributes, reviewed by a human before publish, turns a quarter of work into a fortnight.

  2. Support answers the same questions daily

    A strong case, provided answers come from your actual policies and product data rather than being improvised by a model.

  3. Product data needs categorising and tagging

    A good case. Classification against a defined taxonomy is the kind of judgement models do reliably and people find tedious.

  4. Someone wants a chatbot on the homepage

    Usually the weakest case, and the most commonly requested. It is visible, which is not the same as valuable, and it fails publicly when it is wrong.

A deliberately narrow definition

“AI orchestration” is used to mean a great many things. What we mean by it is specific: connecting a model to your own data and systems so that a repetitive piece of work happens reliably, with a person in the approval path and a measurable error rate.

That excludes a lot of what gets sold under the same heading. We are not proposing an autonomous agent that runs your merchandising, and we would be sceptical of anyone who is.

Why a person stays in the loop

Models are good at producing plausible output and indifferent to whether it is true. In commerce that distinction has a price: a wrong delivery promise is a refund, a wrong specification is a return, a wrong policy answer is a complaint.

So the default is draft-and-approve. Where the measured error rate on a scored sample justifies removing the human step, we remove it deliberately and keep monitoring. Where it does not, the step stays — which is most places, for now.

Where this sits

The useful cases nearly all depend on structured product data, because a model given clean attributes produces specific copy and a model given nothing produces generic copy. Connecting the workflows to your systems is integration work, and if the goal is catalogue legibility for shopping assistants, that is agentic commerce readiness.

Platforms

Where we build.

Start with the boring workflow, not the chatbot

The AI request we receive most often is a customer-facing assistant. It is the hardest to get right, the most damaging when it is wrong, and rarely the best return. The work that reliably pays back is internal: drafting product copy at catalogue scale, classifying and tagging, triaging tickets, flagging data anomalies. Less impressive in a demo, considerably more valuable by the quarter.

Talk through a use case

Scope

What a build includes.

  • Workflows, not demos

    A pipeline that runs on a schedule or a trigger, writes to the systems you already use, and can be turned off without breaking anything.

  • A human approval step

    Generated content is drafted, not published. Where output touches customers, someone signs it off until the error rate justifies otherwise.

  • Grounding in your own data

    Answers drawn from your policies, product attributes and order data, so the system says what is true for you rather than what is generally true.

  • Evaluation before rollout

    A scored sample against known-good answers, so quality is a measurement rather than an impression formed from three lucky examples.

  • Cost and rate-limit controls

    Budget caps, caching and batch processing. Per-token costs are trivial until they run across a full catalogue, then they are not.

  • A documented fallback

    What happens when the provider is down, slow or changes a model. Anything in a customer path needs a defined behaviour for that day.

How it runs

From first call to live.

  1. Find the repetitive work 1–2 weeks

    Where your team spends hours on judgement that is consistent and rule-shaped. That is the shortlist, and it is rarely what was requested.

  2. Prove one case 2–4 weeks

    One workflow, evaluated against a scored sample, with a real cost per run. Enough to decide whether to continue on evidence.

  3. Build and integrate 4–10 weeks

    Into the systems that hold the data and receive the output, with approval steps, logging and the ability to replay a batch.

  4. Monitor Ongoing

    Quality sampling and cost tracking. Model behaviour changes under you, so a workflow that was accurate in March needs rechecking in September.

Frequently asked questions

What is the most useful first AI project in commerce?

Usually product content at scale. If you have thousands of SKUs with structured attributes but thin descriptions, generating reviewed first drafts converts a long manual backlog into a short editorial one. It is measurable, low-risk because a human approves before publish, and it compounds with search and feed quality.

Will AI-written product copy hurt our SEO?

Unreviewed bulk-generated copy can, because it tends to be generic and near-duplicate across similar products. Reviewed drafts built from real attributes are a different thing — the distinguishing factor is whether the content is accurate and specific, not which tool produced the first version. We build the review step in for that reason.

Can you build a customer-facing AI assistant?

Yes, with conditions. It must be grounded in your actual policies and product data, it needs evaluation against a scored sample before launch, and it needs a defined fallback and escalation to a human. What we will not do is connect a general model to your storefront and let it improvise answers about delivery and returns.

How do you control cost?

Batch processing rather than per-request where latency allows, caching repeated work, choosing a model appropriate to the task rather than the largest available, and hard budget caps. Per-token prices look negligible until a job runs across an entire catalogue.

What happens when the model changes?

Output quality shifts, sometimes without announcement. That is why we build evaluation sets at the start and re-run them on a schedule — so a regression is caught by a test rather than by a customer.

Want to talk through AI Orchestration for your store?

Tell us where it hurts. We will tell you honestly whether we are the right people for it.

Talk to an expert Book 30 minutes