You changed the headline. Next month's conversion rate went up.
Good news. But did the headline do it?
The ads may have changed. So might the audience, budget, season or offer. Comparing last month with this month bundles all of that into one number.
If you want to judge the copy, give it a fair comparison.
Keep some visitors on the original
Randomly assign a share of eligible campaign visitors to the original page. Show the rest the campaign version.
Now you're comparing people arriving from the same campaign over the same period. Keep the offer and other page behaviour the same so the copy is the deliberate difference.
That's a randomized holdback. It's also an A/B test.
Calling it a holdback doesn't get around low traffic. It gives you a cleaner comparison with the traffic you have.
Decide what counts as a conversion before starting. A submitted enquiry and a button click answer different questions. Use the same definition for both groups, and check that both are being counted correctly.
Read the counts before the percentage
Here's an illustrative result:
| Version | Visitors | Conversions | Conversion rate |
|---|---|---|---|
| Original | 40 | 2 | 5% |
| Campaign copy | 200 | 14 | 7% |
The campaign version is ahead by two percentage points. That's 40% relative lift.
Sounds strong.
But one more conversion in the original group would put it at 7.5%, ahead of the campaign version.
Nothing about the copy changed. One form submission changed the story.
This is why a lift percentage needs its denominators beside it. “40% better” tells you very little without knowing whether it came from a handful of enquiries or thousands.
Direction is useful. Uncertainty matters.
Read a result in three parts:
- Direction: which version is ahead?
- Size: how big is the observed difference?
- Evidence: how many visitors and conversions sit behind it?
A confidence interval for the difference can help show how much uncertainty remains. With small samples, it may include outcomes where the copy helps, does little or hurts.
Use a method appropriate for small counts if you calculate one. A polished calculator can't make sparse data precise.
An interval that includes zero leaves the direction unresolved at that confidence level. An interval entirely above zero supports an improvement under the test's assumptions. Neither tells you the result will carry unchanged into every future campaign.
Choose your outcome and review point before starting. Repeatedly checking and stopping the first time the numbers look good can make weak evidence look convincing.
You still have a business decision to make
A headline you can reverse in five minutes carries a different risk from a site rebuild or an annual contract.
For a small, reversible edit, you may decide to keep a promising version while collecting more evidence. Describe that honestly: the early result looks encouraging, and it's still uncertain.
For a costly commitment, require stronger evidence.
If the original is ahead, take that seriously too. It may be noise, but the data aren't giving you a reason to declare the new copy a winner.
You can make a practical decision without turning it into a proof claim.
A tiny control group has a cost
Showing most visitors the new copy can feel efficient. It also leaves fewer observations in the original group.
That's the weak point in the example above: forty control visitors and two conversions.
A more balanced split generally estimates the difference more precisely for a fixed total sample. A small holdback trades some of that precision for showing more people the new version.
And if the new version is worse, sending it most of the traffic has a cost of its own.
Choose the split deliberately. Keep visitor assignment consistent so repeat visits don't casually move people between versions.
One page doesn't combine every campaign's evidence
A shared page can reduce maintenance. It doesn't automatically solve measurement.
If each campaign gets different copy, each version's result depends on the visitors who saw it. Pooling everything may hide one version helping while another hurts.
A holdback also answers a narrow question: did the page copy change what visitors did after arriving?
It doesn't establish whether the ads generated sales that would never have happened otherwise.
How Lens approaches it
I build Lens around reviewed campaign copy on an existing page, with a share of campaign visitors held back on the original.
That gives the new copy a comparison. It doesn't promise a decisive answer from a few dozen clicks. If you want a confidence interval, calculate it separately from the underlying counts.
Start with one campaign, one meaningful conversion and working tracking. Look at the counts, keep the uncertainty visible, and give yourself room to learn that the original was better.
Get a Lens setup plan: ten questions and a step-by-step plan for your campaigns. No signup.
