A/B testing in marketing 2026: how not to waste your budget on false insights
80% of A/B tests in marketing produce a false positive result due to insufficient statistical power. How to calculate sample size, when to stop a test, how to prioritize hypotheses. With formulas and real cases.

An A/B test is the only way to know for sure and not guess. Everything else is expert opinion. Opinion is useful for generating hypotheses. The test is useful for checking them. The difference is fundamental.
I tested several hundred hypotheses over eight years - on landing pages, in advertisements, in email newsletters. The findings that surprised me the most: the “Calculate cost” button converted 34% better than “Leave a request” with the same text around it. Removing one field from the form resulted in a 22% increase in conversions. A photo of a real client instead of a stock image - plus 41% CTR.
A marketer's intuition is worth approximately nothing. The test costs a month of work. The test result is worth years of optimization.
1. How to formulate a hypothesis
Incorrect hypothesis: “Let’s change the picture.” This is not a hypothesis. This is an action without justification.
Correct hypothesis: “If you replace a stock photo of an office with a photo of a real team, CTR will increase because people trust real people more than staged photos.”
Hypothesis structure: If [change], then [metric] will change by [expected value] because [mechanism]. This disciplines and prevents you from testing random things.
2. Sample size: when there is enough data
The most common mistake: stopping the test when option B outperformed A by 15% after 200 conversions. This is statistical noise. A practical example: on one project, a test after 3 days showed +22% to the conversion of option B. After a full 14 days, the difference turned out to be −3%. Not significant in the other direction either.
| Basic conversion | Expected growth | Need conversions per option (95%) |
|---|---|---|
| 1% | +20% (up to 1.2%) | ~6300 |
| 2% | +20% (up to 2.4%) | ~3200 |
| 5% | +20% (up to 6%) | ~1300 |
| 2% | +50% (up to 3%) | ~900 |
| 5% | +50% (up to 7.5%) | ~380 |
Conclusion from the table: the lower the base conversion and the smaller the expected effect, the more traffic you need. On landing pages with a conversion rate of 1–2% and sites with traffic of 100–200 visitors per day, the test will take months. It's okay - you just need to plan.
3. What to test first
Prioritization by potential impact: first we test what affects the biggest funnel drop - the weakest transition in the funnel.
If the CTR of an ad is 0.5%, and the norm for the niche is 1.5–2%, we test the ads. If the CTR is normal, but the landing page conversion is 0.8% when the norm is 3–5%, we test the landing page. There is no point in testing CTAs on a landing page if the ads are driving untargeted traffic.
The most highly effective hypotheses in my experience: landing page H1 header, main CTA (text and location), offer (what exactly we offer), number of fields in the form, social proof (type and location).
4. A/B test in advertising accounts
In Direct and VK Ads, ad testing works a little differently than on the website. The algorithms themselves distribute impressions between options, preferring those that give the best CTR. This is not a pure A/B test - it is an optimization of the algorithm.
For a pure test in Direct: create two identical campaigns with different ads, set the same budget and do not touch them at the same time. Compare after 7–14 days using one metric.
VK Ads has a built-in split test function - but remember: the algorithm interferes with traffic distribution. For advertisements this is ok. For landing pages, an external tool is better (Google Optimize or Y.Metrika experiments).
Learn more about testing creatives: Advertising creatives for performance: what works in 2026.
5. Typical A/B testing mistakes
Stopping the test too early. As soon as one option starts to win, the hand reaches out to stop and start the winner. Early stopping gives false results in 30–40% of cases.
Testing several things at the same time. If you replace the image, the headline, and the CTA, it’s not clear what worked. One change at a time.
Seasonality and days of the week are not taken into account. A test run only on weekdays will give a different result than on a full advertising cycle. This is especially critical for B2B (peaks on Mon-Wed) and e-com (peaks on Fri-Sun).
The winner is considered based on CTR, not CPL. The CTR has increased, which means the ad is attracting more clicks. But if the landing page doesn't convert that traffic into leads, it's not a win. The final metric is CPL or ROAS, not CTR.
6. Automation and tools
To test landing pages without a developer: Ya.Metrika experiments (free, for Runet), Google Optimize (closed, but there are analogues - VWO, AB Tasty). For a quick prototype - v0 or Claude Code to create version B of the landing page.
To record hypotheses and results, use a table in Notion. Structure: hypothesis, metric, period, result, conclusion. Without it, in a month you won’t remember what you’ve already tested.
If you want to analyze the testing methodology for your project, write to Telegram @dipustovalov or through form. Starting consultation - 0 ₽.
Related materials: landing page conversion benchmarks, CPA calculator, landing page template, hypotheses for A/B test.