Minimum Detectable Effect: A Practical Guide for Marketers

Learn what minimum detectable effect means, how to calculate it, and why it matters for your A/B test planning and sample size decisions.

Published on 13 min read

Table of contents

The minimum detectable effect is the smallest true change your test can reliably spot given your traffic volume and confidence standards. With 5% significance and 80% power, a common rule of thumb places the threshold at about 2.83 standard errors above zero.

You launch a headline test, split traffic between the original and the variant, and wait. Two weeks later, the conversion rates look almost identical. The team calls it a tie and moves on, but that conclusion may be premature. A flat result can mean the headline made no meaningful difference, or it can mean the experiment was too small and noisy to detect the change you cared about.

That distinction is where minimum detectable effect, or MDE, earns its place in a marketer's planning workflow. It tells you what your design can realistically find before a single visitor sees the experiment. From there, you can decide whether to run the test, collect more traffic, change the target metric, or stop because the likely business impact isn't worth the effort.

Why Your A/B Test Result Might Be Misleading

Your team has a new headline for a paid-traffic landing page. The copy is clearer, the value proposition is more specific, and everyone expects an improvement. You run the original against the new version for two weeks. At the end, the difference is small and the statistical result isn't significant.

A common reaction is, “The headline didn't work.” That's an understandable conclusion, but it isn't the only one. If the experiment could reliably detect only a relatively large change, a smaller real improvement would remain hidden inside normal variation.

A man looking thoughtfully at A/B test conversion rate results displayed on his laptop computer screen.

A flat result isn't proof of no effect

Think of statistical detection as trying to hear a voice in a busy room. A loud announcement is easy to identify. A quiet comment can be real, but background noise may prevent you from distinguishing it from chance variation.

Your conversion rate has the same problem. Visitors differ by device, campaign, geography, intent, returning status, and many other factors. Those differences create variance. When the sample is limited, the signal from a modest change may not stand out clearly enough.

The right question isn't only:

Did the test produce a significant result?

Ask a second question before making a decision:

Was the test sensitive enough to detect the smallest improvement that would matter?

That second question is the practical role of MDE. It prevents a team from treating “not significant” as “nothing happened.”

The missing planning checkpoint

MDE isn't a score assigned to the variant after the test. It's a design quantity calculated in advance from your sample, outcome variability, significance standard, and desired power. The World Bank's DIME guidance on power calculations describes the same relationship: the smallest detectable effect depends on sample size, outcome variance, and test design.

Before launch, write down the smallest change that would justify shipping the headline. Then compare that business threshold with the experiment's MDE. If your design can detect only changes much larger than the improvement you care about, a flat result won't answer the business question.

Teams working with larger ecommerce audiences can also benefit from practical guidance on scaling A/B tests for DTC brands. For a quick pre-launch review of assumptions, allocation, metrics, and stopping rules, use an A/B test sanity check before traffic is split.

What Minimum Detectable Effect Means

Minimum detectable effect is the smallest true effect a study can detect at a specified significance level and power. For a growth team, the practical question is:

If a change this size exists, is the test designed to have a good chance of detecting it?

That question belongs in the planning meeting, before traffic reaches either version. It connects the business threshold for shipping a change with the sample and test duration required to detect it.

Howard S. Bloom formalized MDE in widely used statistical practice in 1995, defining it as the smallest true impact an experiment has a good chance of detecting. His framework connects effect size with the probability of producing a statistically significant estimate at chosen significance and power levels. The original discussion appears in Bloom's evaluation research paper.

The four inputs behind the threshold

MDE becomes clearer when its ingredients are separated:

  1. Sample size is the amount of evidence the test will collect.
  2. Variance is the natural movement in outcomes from visitor to visitor.
  3. Significance level, or alpha, sets the evidence required to call a result statistically significant.
  4. Power is the probability of detecting the effect if it exists.

A standard planning convention uses 5% significance and 80% power. With those settings, MDE is the effect expected to reach statistical significance at the chosen alpha with the chosen probability of detection. The commonly cited relationship of approximately 2.83 standard errors above zero for that combination is discussed in the MDRC explanation of MDE and statistical power.

MDE changes the question you ask

A test report usually presents the observed result: the variant gained, lost, or showed no statistically significant difference. MDE adds the planning question: could this experiment distinguish the improvement the business cares about from ordinary noise?

An effect below the MDE can still exist without producing a statistically significant result under the chosen design. That outcome does not justify claiming a win, nor does it prove that nothing changed. Read it alongside the test's sensitivity and confidence interval before deciding whether to retest, gather more traffic, or stop.

MDE also differs from the lift you expect. It does not predict the variant's eventual result or describe the effect after data collection. It defines the smallest true effect the design can credibly uncover, giving the team a practical basis for setting sample size, estimating duration, and deciding whether the test is worth running.

How Sample Size, Variance, and Confidence Shape MDE

Before launching an experiment, a growth team has to answer three practical questions: how much traffic is required, how long the test may run, and whether the smallest meaningful improvement is detectable. MDE connects those decisions. Change the sample size, the noise in the metric, or the evidence standard, and the test's sensitivity changes too.

Sample size gives you a clearer signal

More observations generally reduce standard error. With less sampling noise, a modest difference has a better chance of standing apart from random variation, so MDE falls as eligible traffic rises, assuming the rest of the design stays constant.

The relationship is not linear. Some additional traffic can help, while making MDE dramatically smaller may require a much larger sample. A team with limited eligible visitors can therefore turn a short experiment into a long wait by choosing an overly precise target.

For a two-proportion z-test, a practical approximation is:

MDE = (z₁₋α/₂ + z₁₋β) √(2p̄(1-p̄)/n)

Here, p̄ is the average baseline proportion, n is the sample size per group, and the two z-values represent the selected significance and power settings. You do not need to calculate this by hand for every test. An A/B test calculator can model alternative traffic, duration, and sensitivity choices before launch.

Variance is the noise in the room

Two landing pages can receive the same number of visitors but produce different levels of measurement noise. One may attract a consistent audience. Another may combine highly qualified branded traffic with low-intent prospecting traffic, making a small lift harder to distinguish.

For conversion experiments, the baseline rate affects the calculation. Revenue and average order value can be noisier still when a few unusually large purchases influence the result. Tighter audience definitions, a stable primary metric, blocking, stratification, or covariate adjustment can improve sensitivity without just waiting for more visitors.

Alpha and power set the evidence standard

A stricter significance requirement raises the evidence threshold for declaring a result. A higher power target asks the test to detect a real effect more consistently. Both choices influence MDE and should match the cost of false positives, false negatives, and the decision that follows.

The MDRC guidance on effects below MDE explains why a commonly used planning combination of 5% significance and 80% power produces the familiar 2.83 standard-error relationship. These settings are planning choices, not laws. Select them before launch, then use the resulting MDE to judge sample requirements, expected duration, and whether the test can support the business decision.

Planning rule: Choose an MDE because the corresponding business effect would justify action, not because a calculator produces a convenient test duration.

When to Accept a Large MDE and When to Aim Lower

A high MDE isn't automatically a flaw. It tells you that the design is sensitive to larger changes and less sensitive to smaller ones. Whether that's acceptable depends on what you're testing and what decision the result will support.

Test situationA larger MDE may be acceptableA lower MDE is more appropriate
Early product conceptYou're screening for a clear directional responseSmall differences will determine whether the concept survives
Major positioning changeOnly a substantial improvement would justify rolloutThe change affects a high-value funnel step
Small audience segmentThe test is exploratory and resource limits are strictThe segment represents a strategic market or customer group
Mature landing pageYou expect only a noticeable improvement to matterIncremental gains can justify implementation work

Match sensitivity to the decision

Start with the smallest effect that would change your action. If a redesign costs significant engineering and review time, a tiny observed movement may not justify deployment even if it eventually becomes statistically distinguishable. If a small improvement compounds across a high-volume acquisition funnel, a broader sample may be worth the wait.

Growth teams often make a costly mistake here. They select an arbitrary target lift because it fits the available traffic, then retrofit the business case around it. The sequence should run the other way. Define the decision threshold first, then ask whether the available design can detect it.

A large MDE can be honest

An experiment with limited eligible traffic may have a large MDE. That isn't a reason to hide the limitation or keep the test running indefinitely. It's a reason to state the scope clearly:

  • What it can detect: A substantial improvement that would justify immediate action.
  • What it can't establish: Whether a small improvement exists.
  • What decision remains open: Whether to collect more data, redesign the test, or accept the uncertainty.

A high MDE is also reasonable when the downside of missing a small effect is low. An exploratory headline test, for example, may be intended to identify a strong winner rather than measure a subtle refinement. A growth-critical checkout change deserves more care because a small effect can still influence a decision with meaningful operational consequences.

A test isn't underpowered in the abstract. It's underpowered relative to the smallest effect you need it to detect.

Worked Examples for Real Marketing Tests

A paid acquisition team is preparing a headline experiment. Before any visitor is assigned to a variant, the team must decide what improvement would justify rollout, how much traffic the test needs, and whether the available launch window can support that decision.

Example one, a headline test

The team will compare the current landing-page headline with a more specific alternative. A 2 percentage-point absolute lift would justify implementation. In practical terms, the variant must produce a conversion rate two percentage points higher than the control baseline.

The planning workflow is:

  1. Define the baseline. Use a stable historical conversion rate from the same page and audience. A baseline from a different traffic mix can distort the sample-size estimate.
  2. Set the meaningful effect. The business threshold is a 2 percentage-point absolute improvement. This is the change worth acting on, not a prediction that the headline will achieve it.
  3. Choose alpha and power. The team uses 5% significance and 80% power, the convention explained in the previous section.
  4. Estimate sample needs. Run a two-proportion power calculation using the baseline rate, target rate, and required sample per group.
  5. Check duration. Compare the required eligible visitors with typical page traffic and the planned allocation between control and variant.

The calculation connects the commercial question to the operating plan. If the required sample exceeds the traffic available during the intended launch window, the team can extend the test, broaden the eligible audience without changing the question, or accept that the design cannot detect the chosen effect reliably.

A target MDE is not a forecast. It is the smallest change the design is being built to detect under the selected statistical standards. A test designed around a 2 percentage-point MDE can still produce a larger, smaller, or zero observed change.

A four-step infographic showing the Real-World Minimum Detectable Effect calculation process with a comparison chart.

For a practical walkthrough of landing-page experimentation, see this guide to a landing page split test. Fix the primary metric and stopping rule before results begin to accumulate, so test duration is not adjusted in response to inconvenient results.

Example two, a lower-traffic segment

Apply the same target lift to a narrow audience, such as visitors from one campaign or returning users on a particular device. The business question remains the same, but the eligible sample is smaller. Holding variance, alpha, and power constant, the design will usually produce a larger MDE.

That changes the launch decision. If the segment cannot provide enough observations during the planned duration, it may detect only a much larger improvement than the team considers meaningful. The team should decide what to do before traffic is split:

  • Go as planned if the segment can detect the effect that would change the decision.
  • Extend the duration if more eligible traffic will arrive without changing the audience definition.
  • Change the design if a more stable metric or broader, pre-specified audience addresses the same business question.
  • Do not launch if the resulting MDE exceeds any plausible improvement worth implementing.

The calculation is a planning filter. It shows whether the experiment can answer the question, how long data collection may take, and what an inconclusive result would mean. This video provides additional visual context for planning an experiment:

Using MDE to Make Better Test Decisions

A disciplined growth team treats MDE as a pre-launch decision tool, not a post-launch excuse. Before traffic is split, record the baseline metric, the smallest meaningful effect, the significance level, the power target, the expected sample, and the decision that follows each outcome.

That short planning record gives stakeholders a shared definition of success. It also exposes weak ideas early. If the required sensitivity demands more traffic than the channel can provide, everyone can discuss the trade-off before copy, design, engineering, and analytics time is committed.

Turn the calculation into a launch gate

Use this sequence in the experiment brief:

  1. Name the decision. State exactly what the team will do if the variant wins, loses, or remains inconclusive.
  2. Define the meaningful effect. Express it in the metric that matters, such as an absolute conversion-rate change or a revenue difference.
  3. Calculate the design MDE. Use the planned sample, variance assumptions, alpha, and power.
  4. Compare the two thresholds. The test is viable when the detectable effect is no larger than the smallest effect worth acting on.
  5. Set the duration and stopping rule. Don't extend a test just because the early result is inconvenient.

If the planned MDE is too large, the answer isn't always “get more traffic.” You might reduce variance with a better-defined audience, choose a more reliable primary outcome, use a more efficient design, or postpone the test until the question can be answered properly.

Report flat results with precision

Avoid telling stakeholders, “The variant had no effect,” when the test wasn't sensitive to the effect size they cared about. Say what the evidence supports:

  • The observed difference wasn't statistically significant.
  • The experiment's MDE was larger than the smallest improvement the team wanted to detect.
  • The result rules out only effects above the design's sensitivity threshold, not every smaller effect.
  • The next decision is to stop, redesign, or collect more evidence.

The GrowthBook power documentation frames MDE as a sensitivity measure for a fixed sample. That makes it especially useful in CRO, where the central question is often whether the test could detect the expected lift at all.

Build the habit across campaigns

MDE belongs in the same planning conversation as audience eligibility, allocation, primary metric, test duration, and implementation cost. It helps a Head of Growth explain why a high-traffic page can support finer-grained experimentation while a narrow segment may need a broader design or a different question.

For landing-page teams, Polish can install a script that reads the page, generates alternative headlines, subheadings, and calls to action, and selects versions for visitors based on factors such as campaign, device, country, and visit status. The same measurement discipline still applies: define the smallest meaningful effect and the decision threshold before interpreting performance.

Use MDE consistently and your team stops asking only whether a test produced a winner. It starts asking whether the experiment was capable of producing a useful answer, which is the standard that protects both budget and learning velocity.


Polish helps you create and evaluate landing-page variations for different visitor contexts, giving your team more ways to test meaningful changes. Visit Polish to see how its visitor-specific page optimization can fit into a disciplined CRO workflow.

  • minimum detectable effect
  • A/B testing
  • statistical power
  • sample size
  • conversion optimization

Share this post

Your website rewrites itself until it converts.

Polish writes new versions of your headlines and CTAs, tests them on your real visitors and keeps the ones that win.

Start free