Blog

Proving It Before You Commit: Testing a Ranking Change on a Slice of Traffic

Sarah Mackinnon
September 3, 2026
Share
Jump to Title

How to design the test, for teams who need evidence before a signature

Every vendor in this category will show you a lift chart. Sponsored clicks up. CTR up. Revenue up. It's an easy chart to produce, because sponsored-click lift is close to the easiest number in retail media to move in the direction you want. It's also one of the least useful numbers for deciding whether to sign a multi-year contract, because almost nobody asks the vendor showing it the one question that would tell you whether to trust it: what measurement design could have proven this didn't work?

If the honest answer is "none, really, the test was built to succeed," that's not evidence. That's a demo.

The fix isn't to distrust every lift chart. It's to design the test yourself, before anyone commits, in a way that could have come back negative. A test that can't fail isn't a test.

TL;DR

Retail media vendors treat measurement as something they supply, a lift number at the end of a pilot, rather than something the retailer designs from the start. Sponsored-click lift is easy to produce and proves almost nothing about what happened to the rest of the page. The fix is a holdout on a slice of traffic, defined before anyone commits, that measures organic conversion alongside sponsored performance rather than sponsored performance alone. Attribution models are mutually exclusive, not additive: a sale counts once, under one method, not stacked across several that each claim it. A coordination layer built to sit alongside an existing stack can run this kind of test on partial traffic without touching the rest of the page, which is what makes the test real rather than hypothetical.

Why Isn't Sponsored-Click Lift Proof of Anything?

Sponsored-click lift answers a narrow question: did more people click on sponsored products than before? It says nothing about where those clicks came from. A sponsored product that outranks a more relevant organic product can generate a real, measurable lift in sponsored clicks while making the page worse overall, because some share of those clicks would have gone to the organic product anyway, and some share of shoppers who would have converted on the organic product bounce instead.

That's the gap a vendor's standard lift report is built to avoid showing you. It measures the ad. It doesn't measure the page. A test that only reports what happened to sponsored performance, without reporting what happened to organic performance on the same pages during the same window, has already decided the answer before the test ran. Andrew Lipsman has made a similar case for why retail media needs to move past a ROAS-only view of performance.

What Does a Defensible Holdout Actually Look Like?

A defensible test holds back a ranking change on a slice of traffic, a percentage of sessions, a set of categories, or a comparable geographic split, while the rest of the site runs as it currently does. The holdout group is your counterfactual: what would have happened anyway, on the same site, in the same window, to the same kind of shopper.

The design choices that make a holdout defensible rather than decorative: the split needs to be large enough and random enough that the two groups are genuinely comparable, not just convenient. The test needs to run long enough to cover normal variation, not just a strong week. And critically, the holdout needs to measure organic conversion and revenue on the same pages where the ranking change is live, not just sponsored performance in isolation. A page-level test tells you what happened to the page. An ad-level test tells you what happened to the ad, and mistakes that for the same thing.

What Does "Radio Button, Not Checkboxes" Actually Mean for Attribution?

This is a framing worth internalizing before you look at any vendor's results: a sale should get attributed to one thing, like a radio button, not credited to several things at once, like a checklist where every box gets ticked. Last-click attribution, incrementality testing, and marketing mix modeling can each produce a defensible number for the same sale, and those numbers are not meant to be added together. A retailer or vendor that reports last-click ROAS and an incrementality lift and an MMM contribution for the same campaign, then implies all three are additive, is overstating performance by definition, because the same converted shopper is being counted under three different logics at once.

The practical version of this discipline: pick the attribution method that matches the question you're actually asking, report it on its own terms, and resist the temptation to stack a second or third method on top to make the number look larger. If a vendor's reporting doesn't make clear which single method produced a given number, that's worth asking about directly before you rely on it.

What Does Incrementality Measurement Actually Test For?

Incrementality measurement asks a more specific question than ROAS does: of the sales that happened during a campaign, how many actually happened because of the campaign, rather than happening anyway? A campaign can post an impressive ROAS while being largely non-incremental, if it's mostly reaching shoppers who were already going to buy. A campaign with a more modest ROAS can be highly incremental if it's genuinely changing shopper behavior. The two numbers answer different questions, and only one of them tells you whether the spend is doing something the business wouldn't have gotten for free.

A holdout, structured the way this piece describes, is the direct way to measure incrementality at the page level: compare converted revenue in the exposed group against the held-out group, on the same pages, over the same window, and the difference is your incremental lift, not just your reported ROAS.

What Does Month Four Actually Look Like?

Most pilots get judged in the window everyone is paying close attention to: the first few weeks, with the champion who pushed for the test still in the room and still checking the dashboard daily. Worth asking before you commit: what does this look like in month four, after the novelty wears off, after the internal champion has moved to a different project or a different company, and after the number has to hold up without anyone actively managing it?

This isn't a reason to distrust every early result. It's a reason to design the test window long enough to include at least one period where performance isn't being actively watched and adjusted, and to name in advance who owns the number after the pilot team's attention naturally moves elsewhere.

What Internal Capability Does This Actually Require?

The honest answer is usually a person, not a feature. Running a genuine holdout, reading the results correctly, and deciding what to do with an ambiguous outcome takes someone on the retailer's side who understands both the statistics and the business context well enough to push back on a vendor's interpretation of their own test. Getting the right people in the room before implementation starts, including whoever owns measurement, is worth treating as its own conversation rather than an afterthought. A retailer without that capability in-house isn't disqualified from running a real test, but should build the cost of getting that expertise, whether through a new hire, a consultant, or a genuinely independent measurement partner, into the decision alongside the vendor's own price.

Who Has the Authority to Call It Off, and Against What Number?

This should be settled before the test starts, not debated after the results come in ambiguous. Name the specific person or committee with authority to end the pilot, and name the specific number or threshold that would trigger that call, in writing, before anyone has a stake in the outcome looking a certain way. A test without a pre-agreed stopping rule tends to run until it produces a number someone likes, which defeats the purpose of testing in the first place.

What Happens to the Rest of Your Stack If the Numbers Don't Hold?

This is the question that determines whether your test was ever real. If a ranking change is implemented as a coordination layer sitting alongside your existing ad server and search engine, rather than as a change to either of those systems directly, the honest answer to "what changes in our underlying stack if this doesn't work out" should be nothing. This is the same principle behind running a pilot alongside an existing stack before committing: the layer runs on a slice of traffic, the holdout produces its result, and if the result doesn't hold up, removing the test changes nothing else about the systems you were already running.

That answer being genuinely "nothing" is the whole point of running the test this way. A pilot that requires touching your core ad server or search infrastructure to set up isn't really reversible, no matter what anyone calls it, because unwinding it means a second migration, not a rollback.

Three Diagnostic Questions

Do you run a holdout on sponsored placement today, or does your current reporting only show you performance for the group that saw the change?

Can you measure organic conversion on the same pages where sponsored inventory changed, or only sponsored performance in isolation?

Who has the authority to call off a test, and against which specific number, decided before the test started rather than after the results came in?

Key Takeaways

  • Sponsored-click lift is easy to produce and measures the ad, not the page. A defensible test measures organic conversion on the same pages during the same window, not sponsored performance alone.
  • A holdout on a slice of traffic is the practical way to measure incrementality at the page level: compare the exposed group against a genuinely comparable held-out group and treat the difference as the real lift.
  • Attribution methods are mutually exclusive, not additive. A sale attributed under last-click, incrementality testing, and marketing mix modeling shouldn't be counted as three separate wins for the same conversion.
  • The internal capability a genuine test requires is usually a person who can read the results correctly, not a feature in a vendor's dashboard.
  • Naming who can call off a test, and against what number, before the test starts is what keeps an ambiguous result from getting explained away after the fact.

Frequently Asked Questions

What is incrementality measurement in retail media? Incrementality measurement isolates the share of sales that a campaign actually caused, as opposed to sales that would have happened anyway. It differs from ROAS, which counts all sales associated with a campaign regardless of whether they were caused by it. A holdout test, comparing an exposed group against a genuinely comparable group that didn't see the change, is the standard way to measure it directly.

What does "radio button, not checkboxes" mean in retail media measurement? It's a framing for how attribution should be reported: a given sale should be attributed under one method, like a radio button selection, rather than credited simultaneously under several methods, like a checklist where every box gets ticked. Last-click attribution, incrementality testing, and marketing mix modeling can each produce a legitimate number for the same sale, but those numbers aren't meant to be added together, since doing so overstates performance by counting the same outcome more than once.

Stay Ahead with Retail Radar

Subscribe for cutting-edge insight into the latest retail media developments and trends

By submitting I accept the Privacy Policy.
Thank you! You are now subscribed to the Pentaleap newsletter.
Oops! Something went wrong while submitting the form.
A mail box
Thank you! You are now subscribed to the Pentaleap newsletter.