Skip to content
In private rollout: onboarded personally, usually live within 24 hours.
Results

Real merchants. No invented logos.

We don’t fake social proof. The case studies below appear only when a merchant has agreed on the record. If the list looks short, it’s because we’d rather have three honest stories than thirty fabricated ones.

  • No borrowed logos
  • No invented metrics
  • No anonymous quotes
In the meantime

Here is how we measure results.

Cervito runs a randomized holdout: a percentage of your visitors never see the widget. Their order rate is the control. Everyone else’s is the test. The lift is a two-proportion z-test, the same statistical method a clinical trial uses. You can see the confidence interval, the sample sizes, and all five attribution models side-by-side, so you can use the one you trust rather than the one that flatters us most.

We tell you this because the first question every merchant asks is: how do I know it’s working? The methodology is the answer. The case studies are evidence the methodology produces real numbers. Both matter.

randomized_holdout · two-proportion z-test · ci + sample_sizes shown

The holdout protocol

Runs every shift

Control

A percentage of your visitors never see the widget. Their order rate is the control.

Test

Everyone else shops with the widget. Their order rate is the test.

The lift is a two-proportion z-test

The same statistical method a clinical trial uses.

Confidence interval · sample sizes · the 5 attribution models, side by side

Last-widgetLinearTime-decayFirst-touchPosition-based

Schematic: the measurement method. No merchant numbers shown.

Not our word for it

The method is already running, live.

The long version

Read the number the way we read it.

What follows is the whole argument: a worked example you can check with a calculator, what the confidence interval should do to a renewal decision, why five attribution models are on screen instead of the one that flatters us, and the questions this method cannot answer.

01

A worked example

Take a fictional store. It sees 40,000 visitors in a calendar month and sets the holdout at 10%. That means 4,000 of those visitors never see the widget at all. They are the control. The other 36,000 shop with it. They are the test. Nobody chooses which group they land in, and neither do we.

At the end of the month the control group has placed 160 orders, an order rate of 4.0%. The test group has placed 1,800, an order rate of 5.0%. The gap between those two rates is the entire result. Everything else on the dashboard is bookkeeping around it.

A fictional month

40,000 visitors · 10% holdout

  • ControlNever saw the widgetVisitors4,000Orders160Order rate4.0%
  • TestShopped with the widgetVisitors36,000Orders1,800Order rate5.0%
Measured lift+25%95% confidence interval: roughly +7% to +46%

Fictional worked example. Not a merchant result, not a benchmark.

One point of order rate does not sound like much. Stated as a lift it is +25%, because 5.0% is a quarter more than 4.0%. Both sentences describe the same fact, and the second is the one every vendor puts on a slide, ourselves included. It is worth knowing that before you compare two vendors’ headlines.

The z-test asks a narrower question: how often would a gap this size show up if the widget did nothing at all and the random split simply fell that way. On these cohort sizes the answer is comfortably under one time in twenty, so the result clears the usual 95% bar. It clears it because the test group is large. The control group, at 160 orders, is the small and expensive half of the experiment, and it is what sets the width of the interval in the next section.

02

What the interval means at renewal

The dashboard does not stop at +25%. It also reports a 95% confidence interval, and on the cohort sizes above that interval runs from roughly +7% to roughly +46%. Read it as a range of plausible truths rather than as decoration on a chart. The honest sentence is not “the widget lifted conversion 25%”. It is “the lift is probably somewhere in that range, and 25% is the middle of it”.

A range that wide is not a flaw in the measurement. It is what 160 control orders buys. Give the same experiment another month and the range narrows, because the control group grew, not because the product improved. That distinction matters when somebody shows you a tightening interval and calls it progress.

When the renewal invoice arrives, the number to work with is the bottom of the range, not the middle. Take the lower bound, apply it to your own order volume and your own gross margin, and see whether it covers what we charge. If +7% covers it, the decision is easy and you never needed the optimistic half of your own data. If only +25% covers it, you are renewing on the flattering end of a range, and the better move is almost always to run another month and look again. We would rather lose a renewal to that arithmetic than win one against it.

Two traps sit beside this. The first is that significant does not mean large: a lift can clear the 95% bar and still be worth less than what you pay for it, particularly on a small average basket, and the statistics will not warn you about that. The second is that a range which still includes zero is not a verdict against the product. It usually means there is not enough traffic yet to tell, which is a different sentence, and we would rather write the true one than round it up into a headline.

03

Five models, not the flattering one

The holdout answers how much extra revenue exists. It does not answer which touch deserves the credit for it. Those are separate questions, and the second one has no single correct answer. That is why the dashboard carries five attribution models side by side instead of quietly picking one for you.

An attribution model is just a rule for splitting credit across a shopper’s journey. All five run over the same 30-day window on the same recorded path, so nothing about the underlying data changes between them. Only the rule changes.

Last-widget
All of the credit to the last widget touch before the order.
Linear
The credit split evenly across every touch in the journey.
Time-decay
Recent touches weighted more heavily than early ones.
First-touch
All of the credit to the conversation that opened the journey.
Position-based
Most of the credit to the first and last touch, the rest spread across the middle.

A vendor showing you one model has already made a choice on your behalf about which story to tell, and in our experience it is not the conservative one. We show all of them at once because the disagreement between them is itself information. When last-widget and first-touch land close together, she is doing the whole job inside one conversation. When first-touch runs much higher, she is opening journeys that close somewhere else, which is a real contribution and an easy one to miss. When they diverge sharply in either direction, that is a reason to look at the journeys before you draw a conclusion.

One thing not to do with them: the holdout lift and the attributed revenue are two views of the same money, not two amounts of it. Adding them together double counts, and a dashboard that invites you to do it is telling you something about the vendor.

04

What this cannot tell you

A method is only worth trusting if it is specific about its own edges. These are ours, and every one of them is checkable on your own dashboard rather than taken on faith.

If you never switch it on
The holdout is opt-in and you configure it, because it means deliberately withholding the assistant from a slice of your own traffic. A store that never enables it has no control group. In that case the dashboard shows engagement, labelled as engagement, and not as lift.
On low traffic
Below a certain volume the interval stays wide enough to be useless, and no amount of dashboard design fixes that. The honest answer is to keep collecting, not to squint at the middle of the range.
If you reconfigure it mid-flight
Change the holdout percentage in week three and you have two short experiments rather than one longer one. The comparison restarts.
If you change several things at once
When a campaign starts, the coaching notes are rewritten, and the holdout moves in the same fortnight, the lift is still real and its cause is no longer attributable. One change at a time is the price of a clean answer.
Off the storefront
Email and SMS exist in the attribution framework but are not wired into live measurement today, so revenue that closes entirely off-site sits outside the experiment.
Beyond the window
Touches count for 30 days, a window chosen deliberately for gifting and considered purchases where the decision takes weeks. Repeat purchase behaviour, returns, and lifetime value are past its edge, and we do not model them.
Sliced after the fact
Cutting a finished result by device, country, or product category produces intervals the sample was never sized to support. Read those as hints worth a follow-up, not as findings.
Against another vendor
No holdout can tell you what a different assistant would have done on your store. The only honest comparison is to run one, then run the other.

None of that makes the measurement weak. A number with published limits is worth more than a number without them, because you can tell when to act on it and when to wait. That is the whole reason this page exists before any case study does.

opt-in holdout · two-proportion z-test · 30-day window · 5 models, no default

In progress

First case studies coming soon.

We’re working with our first cohort of merchants to publish their numbers. If you’re already running Cervito and want to share what you’ve seen, we’d love to feature you.

The case-study ledger

On the record, or not at all

  1. 001The first cohort, being measuredPublished only with the merchant’s sign-offIn progress
  2. 002A merchant story, on the recordPublished only with the merchant’s sign-offReserved
  3. 003A merchant story, on the recordPublished only with the merchant’s sign-offReserved
Your next hire

Ready to put a real associate on your storefront?

Book a 30-minute demo with the founder over Google Meet and watch Sarah work your own catalog. We onboard you personally, usually live within 24 hours. Not ready for a call? Request an invite.

Private rollout · one-click uninstall if it doesn’t earn its keep