Real merchants. No invented logos.
We don’t fake social proof. The case studies below appear only when a merchant has agreed on the record. If the list looks short, it’s because we’d rather have three honest stories than thirty fabricated ones.
- No borrowed logos
- No invented metrics
- No anonymous quotes
Here is how we measure results.
Cervito runs a randomized holdout: a percentage of your visitors never see the widget. Their order rate is the control. Everyone else’s is the test. The lift is a two-proportion z-test, the same statistical method a clinical trial uses. You can see the confidence interval, the sample sizes, and all five attribution models side-by-side, so you can use the one you trust rather than the one that flatters us most.
We tell you this because the first question every merchant asks is: how do I know it’s working? The methodology is the answer. The case studies are evidence the methodology produces real numbers. Both matter.
randomized_holdout · two-proportion z-test · ci + sample_sizes shown
The holdout protocol
Runs every shiftControl
A percentage of your visitors never see the widget. Their order rate is the control.
Test
Everyone else shops with the widget. Their order rate is the test.
The lift is a two-proportion z-test
The same statistical method a clinical trial uses.
Confidence interval · sample sizes · the 5 attribution models, side by side
Schematic: the measurement method. No merchant numbers shown.
The method is already running, live.
Three demo storefronts, three catalogs with nothing in common, one associate. No login, and the same widget build a merchant installs. Open one and ask for something the catalog does not carry, then watch her say so rather than invent a product to fill the silence. That refusal is the part of this method you can check in about a minute, without a spreadsheet.
- Pawfield Pet Co.Independent pet shopAn independent pet shop. Ask for a bed for a senior dog, grain-free food for a sensitive cat, or what you need to start an aquarium.Openthe Pawfield Pet Co. demo store (opens in a new tab)
- Voltline ElectronicsConsumer electronicsA consumer-electronics store. Compare noise-cancelling headphones, spec a work-from-home desk, or find a tech gift under €100.Openthe Voltline Electronics demo store (opens in a new tab)
- Northbrook Coffee Co.Specialty coffeeA specialty coffee roaster. Bundle a pour-over starter set, ask for beans that suit the way you brew, or find a gift for someone particular about coffee.Openthe Northbrook Coffee Co. demo store (opens in a new tab)
All of them, with what to ask and what she is allowed to refuse, are on the demo page.
Read the number the way we read it.
What follows is the whole argument: a worked example you can check with a calculator, what the confidence interval should do to a renewal decision, why five attribution models are on screen instead of the one that flatters us, and the questions this method cannot answer.
A worked example
Take a fictional store. It sees 40,000 visitors in a calendar month and sets the holdout at 10%. That means 4,000 of those visitors never see the widget at all. They are the control. The other 36,000 shop with it. They are the test. Nobody chooses which group they land in, and neither do we.
At the end of the month the control group has placed 160 orders, an order rate of 4.0%. The test group has placed 1,800, an order rate of 5.0%. The gap between those two rates is the entire result. Everything else on the dashboard is bookkeeping around it.
A fictional month
40,000 visitors · 10% holdout
- ControlNever saw the widgetVisitors4,000Orders160Order rate4.0%
- TestShopped with the widgetVisitors36,000Orders1,800Order rate5.0%
Fictional worked example. Not a merchant result, not a benchmark.
One point of order rate does not sound like much. Stated as a lift it is +25%, because 5.0% is a quarter more than 4.0%. Both sentences describe the same fact, and the second is the one every vendor puts on a slide, ourselves included. It is worth knowing that before you compare two vendors’ headlines.
The z-test asks a narrower question: how often would a gap this size show up if the widget did nothing at all and the random split simply fell that way. On these cohort sizes the answer is comfortably under one time in twenty, so the result clears the usual 95% bar. It clears it because the test group is large. The control group, at 160 orders, is the small and expensive half of the experiment, and it is what sets the width of the interval in the next section.
What the interval means at renewal
The dashboard does not stop at +25%. It also reports a 95% confidence interval, and on the cohort sizes above that interval runs from roughly +7% to roughly +46%. Read it as a range of plausible truths rather than as decoration on a chart. The honest sentence is not “the widget lifted conversion 25%”. It is “the lift is probably somewhere in that range, and 25% is the middle of it”.
A range that wide is not a flaw in the measurement. It is what 160 control orders buys. Give the same experiment another month and the range narrows, because the control group grew, not because the product improved. That distinction matters when somebody shows you a tightening interval and calls it progress.
When the renewal invoice arrives, the number to work with is the bottom of the range, not the middle. Take the lower bound, apply it to your own order volume and your own gross margin, and see whether it covers what we charge. If +7% covers it, the decision is easy and you never needed the optimistic half of your own data. If only +25% covers it, you are renewing on the flattering end of a range, and the better move is almost always to run another month and look again. We would rather lose a renewal to that arithmetic than win one against it.
Two traps sit beside this. The first is that significant does not mean large: a lift can clear the 95% bar and still be worth less than what you pay for it, particularly on a small average basket, and the statistics will not warn you about that. The second is that a range which still includes zero is not a verdict against the product. It usually means there is not enough traffic yet to tell, which is a different sentence, and we would rather write the true one than round it up into a headline.
Five models, not the flattering one
The holdout answers how much extra revenue exists. It does not answer which touch deserves the credit for it. Those are separate questions, and the second one has no single correct answer. That is why the dashboard carries five attribution models side by side instead of quietly picking one for you.
An attribution model is just a rule for splitting credit across a shopper’s journey. All five run over the same 30-day window on the same recorded path, so nothing about the underlying data changes between them. Only the rule changes.
- Last-widget
- All of the credit to the last widget touch before the order.
- Linear
- The credit split evenly across every touch in the journey.
- Time-decay
- Recent touches weighted more heavily than early ones.
- First-touch
- All of the credit to the conversation that opened the journey.
- Position-based
- Most of the credit to the first and last touch, the rest spread across the middle.
A vendor showing you one model has already made a choice on your behalf about which story to tell, and in our experience it is not the conservative one. We show all of them at once because the disagreement between them is itself information. When last-widget and first-touch land close together, she is doing the whole job inside one conversation. When first-touch runs much higher, she is opening journeys that close somewhere else, which is a real contribution and an easy one to miss. When they diverge sharply in either direction, that is a reason to look at the journeys before you draw a conclusion.
One thing not to do with them: the holdout lift and the attributed revenue are two views of the same money, not two amounts of it. Adding them together double counts, and a dashboard that invites you to do it is telling you something about the vendor.
What this cannot tell you
A method is only worth trusting if it is specific about its own edges. These are ours, and every one of them is checkable on your own dashboard rather than taken on faith.
- If you never switch it on
- The holdout is opt-in and you configure it, because it means deliberately withholding the assistant from a slice of your own traffic. A store that never enables it has no control group. In that case the dashboard shows engagement, labelled as engagement, and not as lift.
- On low traffic
- Below a certain volume the interval stays wide enough to be useless, and no amount of dashboard design fixes that. The honest answer is to keep collecting, not to squint at the middle of the range.
- If you reconfigure it mid-flight
- Change the holdout percentage in week three and you have two short experiments rather than one longer one. The comparison restarts.
- If you change several things at once
- When a campaign starts, the coaching notes are rewritten, and the holdout moves in the same fortnight, the lift is still real and its cause is no longer attributable. One change at a time is the price of a clean answer.
- Off the storefront
- Email and SMS exist in the attribution framework but are not wired into live measurement today, so revenue that closes entirely off-site sits outside the experiment.
- Beyond the window
- Touches count for 30 days, a window chosen deliberately for gifting and considered purchases where the decision takes weeks. Repeat purchase behaviour, returns, and lifetime value are past its edge, and we do not model them.
- Sliced after the fact
- Cutting a finished result by device, country, or product category produces intervals the sample was never sized to support. Read those as hints worth a follow-up, not as findings.
- Against another vendor
- No holdout can tell you what a different assistant would have done on your store. The only honest comparison is to run one, then run the other.
None of that makes the measurement weak. A number with published limits is worth more than a number without them, because you can tell when to act on it and when to wait. That is the whole reason this page exists before any case study does.
opt-in holdout · two-proportion z-test · 30-day window · 5 models, no default
First case studies coming soon.
We’re working with our first cohort of merchants to publish their numbers. If you’re already running Cervito and want to share what you’ve seen, we’d love to feature you.
The case-study ledger
On the record, or not at all
- 001The first cohort, being measuredPublished only with the merchant’s sign-offIn progress
- 002A merchant story, on the recordPublished only with the merchant’s sign-offReserved
- 003A merchant story, on the recordPublished only with the merchant’s sign-offReserved
Ready to put a real associate on your storefront?
Book a 30-minute demo with the founder over Google Meet and watch Sarah work your own catalog. We onboard you personally, usually live within 24 hours. Not ready for a call? Request an invite.
Private rollout · one-click uninstall if it doesn’t earn its keep