Back to Blog
EngineeringAug 4, 2026· 7 min read

Engaged Shoppers Convert Better. That Is Not a Lift Number.

Every AI shopping assistant on the market quotes a conversion multiple. Nearly all of those numbers compare shoppers who engaged to shoppers who did not, which measures who opens a chat window more than it measures what the chat window did. We have published a version of this number too. Here is what it can and cannot support.

The category has settled on a stat shape. It goes something like: shoppers who used the AI assistant converted at 12%, against 3% for everyone else. Four times better. Sometimes it is five times, sometimes twelve.

The number is usually real. The inference drawn from it is usually not.

What that comparison actually measures

Think about who opens a chat widget on a product page.

Someone comparing two jackets they are close to buying. Someone who needs to know whether the size runs small before they commit. Someone checking the return window because they have already decided and are managing the downside.

Now think about who does not. Someone who bounced in four seconds. Someone who landed from a paid social ad and is browsing. Someone on their eleventh visit who already knows the answer.

The engaged group was going to convert better than the unengaged group if the widget had never been installed. Intent causes engagement and intent causes purchase. Comparing the two groups measures the intent gap plus whatever the assistant contributed, with no way to separate them.

The question is not whether engaged shoppers convert better. They do. The question is what would have happened to those same shoppers with no assistant on the page.

That counterfactual is not observable in your analytics. It has to be constructed.

The design that answers it: an incrementality test

A holdout, which performance marketers know as an incrementality test. Randomly withhold the assistant from a slice of visitors, then compare the two randomized populations rather than the two self-selected ones.

Four things have to be right, and each of them is a place we have seen this go wrong.

Randomize at the visitor, and keep them there

Assignment has to be stable across sessions and devices as far as your identity resolution can carry it. Randomizing per session lets the same person land in both arms across a multi-visit purchase path, which is most considered purchases.

The subtle failure here is worse than the obvious one. If you change the split, say from 50/50 to 90/10 because the results look good, a naive bucketing implementation re-stamps returning visitors into new cohorts. Your history is now a blend of people who saw the assistant, then did not, then did. We have hit exactly this. The experiment does not error. It just quietly stops meaning anything, and the readout looks plausible the whole time.

Measure revenue per visitor, not conversion rate of engagers

This is the one that matters most, and it is entirely about denominators.

The treatment arm is everyone who could have used the assistant, including the large majority who ignored it. If you narrow to engagers you have thrown away the randomization and rebuilt the original selection effect inside your experiment.

Revenue per visitor across the whole arm also catches the effects a conversion rate misses. An assistant that raises average order value while leaving conversion flat is working. An assistant that converts more people into smaller orders may not be. Both are invisible in a conversion-rate readout.

Size it before you run it, and accept the answer

Ecommerce conversion rates are low and order values are skewed, which is a bad combination for statistical power. Detecting a few percent difference in revenue per visitor takes far more traffic than most people assume, typically tens of thousands of visitors per arm.

The honest consequence: plenty of stores cannot detect the effect size they care about in a reasonable window. That is a real finding about the measurement, not a reason to run the test anyway and report whatever came out. A 90/10 split makes this worse, because the small arm sets your power.

Decide the window and the metric before you look. Revenue-per-visitor differences wander early and settle late, and a test you can stop whenever it looks good is not a test.

Know your attribution ceiling and state it

Linking a conversation to an order means carrying an identifier from a chat session through to a purchase event. That chain breaks in ordinary ways: a shopper switches from phone to laptop, storage gets cleared, a privacy setting drops the identifier, the purchase arrives by webhook without the client-side context attached.

So there is a ceiling on the share of influenced orders you can attribute, and it is well under 100% on real storefronts. Anyone claiming complete AI revenue attribution is either not looking or not saying.

Two things follow. Publish your join rate alongside your attributed revenue, because a number without its coverage is unreadable. And prefer the holdout for the causal question, since a randomized comparison of totals does not depend on joining individual orders to individual conversations at all. The identity chain is a reporting problem. It should not be your evidence base.

What we do and do not claim

At Kinect we run holdouts on live storefronts. We are not publishing a lift number from them today, because the ones with enough traffic to be powered have not run long enough, and the ones that have run long enough do not have the traffic. When that changes we will publish the result, including if it is smaller than what the category advertises.

In the meantime, here is the discipline we hold ourselves to, which is also written into our measurement methodology.

ClaimWhat it rests onWhat it is not
Shoppers who engage convert at 6–21%Observed behavior of the engaged cohortStore-wide lift caused by the assistant
Engaged shoppers convert at 5–8× the site averageCross-store comparison of the same two cohortsA causal multiple
3–6% more revenue, measured against their own baselinesBefore-and-after against the store's own historyA randomized causal estimate
A guaranteed lift percentageNothingSomething we will ever publish

The engaged-cohort numbers are labeled as engaged-cohort numbers every single time they appear on this site. That is a deliberate constraint and it costs us in sales conversations, because the unqualified version of the same statistic is a better headline.

The baseline comparison is more useful than a cohort split and still not a randomized estimate. Seasonality, promotions, traffic mix, and every other change you shipped that quarter are all inside it. We say "measured against their own baselines" rather than "caused by" for that reason.

Why nobody has published one

As far as we can tell, no vendor in this category has published a completed holdout result for an on-site AI shopping assistant. Not us either, yet.

The reason is not mysterious. A holdout can come back small. The engaged-cohort number never does, it requires no experimental infrastructure, and it is available on day one of a pilot. Given a choice between a statistic that always flatters and one that might not, the market has made the obvious choice.

The broader pattern is worth knowing if you buy any performance software: across advertising, where incrementality testing is far more mature, measured causal effects routinely come in well below platform-reported ones. There is no particular reason to expect AI assistants to be the exception to a rule that holds across every self-reported attribution system anyone has bothered to test.

Which is the actual argument for running the test. Not because the answer will be flattering, but because a number you can defend in a board meeting is worth more than a bigger one you cannot.

What to ask a vendor

  • Is this number a comparison of engagers to non-engagers? If yes, it is a selection effect and everyone in the room should say so out loud.
  • What is the denominator? Revenue per visitor across the whole arm, or conversion rate among people who engaged?
  • Was assignment randomized, and was it stable across sessions? Ask specifically what happens to returning visitors when the split changes.
  • What share of orders can you actually join to a session? If they do not know, the attributed revenue figure has no error bar.
  • Was the window and the metric fixed before the test ran?
  • Can you show me a result that came back flat? A vendor whose every measurement is positive is running a marketing function, not a measurement one.

Ask us the same questions. If you want to see the methodology we hold ourselves to, it is on how we measure. If you would rather have the conversation directly, book a demo.

Frequently asked questions

What is a selection effect in this context?

Shoppers who choose to open an AI assistant already have higher purchase intent than shoppers who do not. Comparing the two groups measures that pre-existing intent gap plus whatever the assistant contributed, with no way to separate the two. The comparison is real but it does not isolate the assistant's effect.

Why measure revenue per visitor instead of conversion rate?

Because the treatment arm includes everyone who could have used the assistant, not only those who did. Narrowing to engagers discards the randomization and rebuilds the original selection effect inside the experiment. Revenue per visitor also captures order-value changes that a conversion rate misses entirely.

How much traffic does a holdout need?

More than most people expect. Ecommerce conversion rates are low and order values are skewed, so detecting a few percent difference in revenue per visitor typically requires tens of thousands of visitors per arm. With an uneven split such as 90/10, the small arm determines your power.

Can AI-influenced revenue be fully attributed?

No. Linking a conversation to an order depends on carrying an identifier across sessions and devices, and that chain breaks for ordinary reasons including device switching, cleared storage, privacy settings, and server-side purchase events arriving without client context. Any credible attributed-revenue figure should be published alongside its join rate.

What is incrementality testing for an AI shopping assistant?

Incrementality testing holds out a random slice of visitors who never see the assistant, then compares revenue per visitor between the holdout and everyone else. The gap is the assistant's causal contribution. It is the same design marketers use to measure ad lift, and it is the only way to turn an engaged-cohort stat into a real lift number.

The Kinect essays

New essays, in your inbox.

We write about AI shopping, intent, and what happens to commerce when every surface can answer questions. No cadence, no spam — just the next piece when it ships.