Marketing Experiment Design with (un)Common Logic
Marketers talk a lot about testing, yet the gap between a neat A/B idea and a decision you would stake budget on can be wide. I have sat in rooms where a team celebrated a two percent lift that later vanished when the promo calendar changed, and in other rooms where a null test quietly saved seven figures because it revealed an offer that looked attractive in a dashboard but carried a hidden margin reef. Thoughtful experiment design is the bridge between curiosity and conviction. It is also a practical craft. You earn reliability not through complexity for its own sake, but by asking disciplined questions in the language of the business and by designing around the actual physics of the channels you use.

I call that blend of practicality and rigor an (un)Common Logic. It is common because the principles are no mystery, uncommon because they are applied steadily, even when there is pressure to skip steps. Whether you work at a scrappy startup or inside a mature growth engine, the mindset is the same: define the decision, architect the test to isolate the cause, measure what truly matters, and adjust for reality without fooling yourself.
Start from a decision, not a hypothesisGood experiments begin with a decision you are ready to make if the evidence is clear. That discipline cleans up every downstream choice. If the real decision is whether to roll out a new onboarding flow to all new users next quarter, write it plainly. The hypothesis is only a means to that end.
Tie the decision to a target metric the business values. I like to formalize this with a simple statement that fits on one line: We will ship variant B to 100 percent of new signups if it increases 8-week paid conversion rate by at least 5 percent, with no more than a 3 percent drop in average order value. That single sentence nails down the primary metric, puts a line in the sand for minimum practical impact, and introduces a guardrail. It makes sample size and duration solvable. It also inoculates you against the common trap of celebrating statistically significant but commercially irrelevant bumps.
Be specific about the unit of analysis. If the metric is downstream and accumulates over weeks, you probably need user-level randomization, not session-level. If you cannot reliably identify users due to privacy changes, you may need geo-level or time-based designs.
Choose metrics you can defend on a hard dayPrimary metrics should reflect value creation, not proxy engagement. When testing a landing page, click-through rate can be a leading indicator, but revenue per visitor, qualified lead rate, or paid conversion rate is what funds payroll. I have seen teams optimize an email on open rate only to learn that the catchy subject line inflated opens and depressed clicks from their best customers. If you must use a leading metric to shorten test cycles, at least validate its relationship to the business outcome first. Quantify that relationship historically across several campaigns and compute the elasticity. If a 1 point lift in click-through has produced anywhere from a 0.3 to 0.8 point lift in conversions depending on seasonality, build that uncertainty into your expected value.
Guardrails are not decoration. They protect margin, inventory health, unsubscribe rates, page performance, and brand safety. When we tested a more aggressive discount rail on a retail homepage, the primary metric, revenue per session, looked great in week one. The guardrail metric, coupon redemption among full-price buyers over the following two weeks, flashed red. Without that guardrail, we would have taught the most valuable segment to wait for deals, and we would have paid for it for months.
Pre-period adjustments earn their keep too. If you can measure a stable pre-experiment baseline at the unit level, you can use it to reduce variance. Methods like CUPED, which regress outcomes on pre-period means to adjust post-period outcomes, often cut variance by 10 to 40 percent depending on the stability of your customers’ behavior. That is less sample size, or more precision for the same traffic.
Power, precision, and minimum detectable effects you can explain to financeThe right sample size is not a math trophy, it is a commitment to detect only those effects worth acting on. Choose the minimum detectable effect by working backward from the economics of the decision. If shipping the variant would require engineering effort worth 100 person-hours and a promotional budget shift of 150,000 dollars, a 0.5 percent lift in conversion is not worth it unless you have unusually massive volume. A 3 to 5 percent lift might be. Quantify the threshold, then size for that.

A concrete path: fix Type I error at 5 percent, Type II error at 20 percent for 80 percent power, and use a conservative estimate for baseline conversion. If baseline paid conversion is 8 percent and you care about a 5 percent relative lift, that is an absolute increase to 8.4 percent. Plugging those into a two-proportion power calculator yields roughly 64,000 users per group. If your signups run 8,000 per day, the test will need at least eight days plus a buffer for weekday effects. If you can apply a variance reduction strategy that halves variance, you can cut duration by about 30 percent. Do not promise a two-day win unless you can justify the assumptions. Leaders can handle a steady cadence better than missed mini deadlines.
Sequential looks are tempting because everyone wants early reads. They are fine if you use a proper alpha spending plan or a Bayesian sequential approach with predefined decision thresholds. They are dangerous if you peek daily and declare victory on a Friday afternoon because the chart looks pretty. I have watched uplift drift shrink over two weeks due to coupon stacking and delayed churn. Build stopping rules upfront. If you choose a Bayesian approach, define the decision in terms of the posterior probability that the lift exceeds the minimum practical effect, not just that it is above zero.
Randomization where interference will not corrupt itRandomizing at the wrong layer is the fastest way to learn nothing. Digital marketing gives you choices: cookie-level, user-level, session-level, account-level, geo-level, and time-based switchbacks. Each has interference risks and practicality constraints.
User-level randomization is the first choice for product and website tests where identification is stable. It avoids the duplicates and cross-contamination that plague cookie-based approaches. Post-iOS privacy changes have made stable identification in ads and mobile trickier, so you often need to move up a layer.
Geo-experiments work surprisingly well when the outcome is sales by region or store. Think state-level or DMA-level splits. Use 60 to 200 geos if possible, balance them on pre-period outcomes with synthetic control or matched pairs, and run long enough to wash out weekly cyclicality. When we ran a geo-lift test for a national brand on connected TV spend, we used 96 DMAs, blocked them into 48 matched pairs on trailing four-week revenue and audience mix, and randomized within pairs. The result was precise enough to detect a 4 percent lift on a two-week run, something a naive aggregate before-after would have missed by a mile.
Switchback tests shine when your treatment affects the environment, not the user. Ad auctions and delivery algorithms are a good example. If your treatment is a different bidding strategy, toggling it on and off by hour or day while holding everything else constant helps isolate the effect without persistent cross-arm spillovers. The cadence needs to be slower than the system’s memory. If a platform’s learning resets over roughly 48 hours, do not switch every 6 hours. Use 2 to 3 day blocks.
The messy reality of ad platform experimentsPlatforms offer their own testing tools, each with quirks. Facebook’s conversion lift studies and Google’s geo experiments can be powerful, but you need to read the fine print.
With Facebook lift, the holdout is created by withholding delivery to a randomized subset. That makes incrementality estimates cleaner than in-account A/Bs, which mostly compare creatives within the same auction environment. But it also means your campaign structure, budget caps, and learning phase behavior will differ https://privatebin.net/?dbae85c9699dc390#6MrQfDADTgLMEtRhDsknHLpd729W3zCiPXTmqC36Vg8z with and without the holdout. Monitor delivery so that the test arm does not hit artificial constraints. Expect some ghost ad measurement noise for small accounts. Prepare stakeholders for the possibility that an exciting creative within-account wins on cost per result but shows no incremental lift when measured against a holdout. That paradox is common when a creative simply steals from your own other ads.
With Google’s geo experiments, match geos on pre-experiment sales, traffic, and audience composition. Spend must be high enough within treatment geos to generate measurable signal. If you split DMAs and then throttle spend uniformly, you risk under-delivering in your highest potential areas. A better move is to reallocate budget proportionally within treatment geos to sustain impression share. You will get cries of bias. The answer is to use pre-registered reallocation rules and symmetric handling across treatment and control.
Attribution fights will flare. Multi-touch last-click dashboards often diverge from lift estimates because they are answering different questions. When a lift test says your branded search campaign is 90 percent cannibalistic, the natural response is disbelief. Lean on math and transparency. Show how the holdout behaves, show the confidence intervals, and run confirmation tests that move budget out of the cannibal and into a prospecting campaign. The blended return is what matters at planning time.
Duration, seasonality, and the shape of behaviorDay of week effects matter more than people admit. If your DTC site’s weekend traffic converts 1.5 times weekday, a 7-day test is the rock bottom minimum. Better, run two full weeks to capture two weekends and reduce the chance of an odd Monday email blast skewing results. Longer cycles are essential for behavior with lags. If your subscription takes two weeks to activate on average and churn often occurs around week six, a 10-day test on trial signups tells you little about revenue. Define observation windows aligned to behavior, then decide whether to analyze early indicators with a validated mapping to downstream value.
When you test prices or promotions, remember customers learn. The first week of a new promo may pull forward demand, then the effect decays. I once watched a three-week test of a 20 percent off banner show a 12 percent revenue lift in week one that settled to 3 percent net by week three. If we had ended early, we would have captured the initial spike and shipped a policy that eroded margin for months. Use time-series plots, not just aggregates, and model trend plus level change. If the effect is not stable after two cycles, extend or plan a second-phase test with a longer horizon.
Instrumentation and the curse of missing conversionsYour experiment is only as good as your events. I have had beautiful randomization undone by a single untagged pathway. Check that all eligible users can enter both arms, that conversion events are de-duplicated across platforms, and that server-side and client-side events reconcile within a small tolerance. For paid media, align conversion windows with the product reality. A 1-day view-through credit on a 14-day decision cycle will warp creative tests toward clickbait. If you cannot change platform windows, at least analyze exported logs with your own windows.
Conversion lags are not just an annoyance. They change how you stop. If 40 percent of conversions land after day 7, do not lock the test at day 8 and declare winners on partial data that will backfill differently across arms. Either wait for the majority of conversions to clear or use survival analysis and lag-aware models to estimate final outcomes. Keep a concordance check: do late conversions land proportionally across arms, or is one arm systematically late due to funnel friction?
The skeletal checklist that prevents regretWhen time is tight, a small checklist protects you from the most expensive mistakes. Keep it short enough that people actually use it.
Name the decision, primary metric, guardrails, and minimum practical effect in a single crisp sentence everyone agrees on. Choose the randomization unit that matches the interference risk, then write down why not the others. Size the sample for power at the minimum practical effect, and write the stop rules so you are not improvising later. Pre-commit the analysis plan, including any variance reduction, segment cuts, and how you will handle lags. Define how the result maps to an action, including rollout plan, monitoring, and fallbacks if the effect decays.Tape that list on the wall. If a test proposal cannot pass it in 15 minutes, postpone, then fix the gaps.
Analysis plans you can defend without a statistics degreeFor binary outcomes like conversion, difference in means with robust standard errors gets you far, especially with user-level randomization. If your pre-period baselines are strong predictors, apply pre-period adjustment via covariance or CUPED. For count outcomes with heavy tails, such as revenue per user, use trimmed means or a winsorized mean alongside a nonparametric bootstrap to estimate uncertainty. You will sleep better when one outlier does not flip your sign.
Segment carefully. Pre-register two or three slices that reflect meaningful strategy, like new versus returning, paid versus organic, mobile versus desktop. Do not dredge 20 cuts until you find a green box. If you must explore, label it exploratory and run a follow-up confirmation test.
For geo or time-based designs, synthetic control and difference-in-differences are your friends. Build a model to predict the treated unit from a weighted blend of controls in the pre-period, then compare observed to predicted in the post period. Check parallel trends visually. If trends diverge before the treatment, no method saves you. Redesign.
Avoid the allure of uplift modeling unless you have the traffic and infrastructure to deliver targeted treatments at the unit level. Many uplift models fit to noise and then drive dangerous heterogeneity claims. If you do try them, run shadow assignments and holdouts to quantify the real incremental gain versus a simple segment rule.
Decisions under uncertainty, not just p-valuesExecutives remember actions, not p-values. Translate results into expected value with uncertainty. If variant B has a 75 percent posterior probability of delivering at least a 4 percent lift, and your minimum practical effect is 5 percent, what should you do? Sometimes shipping is still right if the downside cost is small and the monitoring plan is strong. Sometimes you hold back because the rollout risk dwarfs the upside.
Frame trade-offs explicitly. If an email subject test shows a 3 percent click lift but a small rise in unsubscribes among high lifetime value customers, show the blended cohort value over six months. A concise decision matrix helps: ship now with guardrails, run a second test focused on the sensitive segment, or table the idea in favor of a bigger lever. That is the heartbeat of (un)Common Logic, the willingness to weigh imperfect signals against real costs.
When a test “does not work,” squeeze value from it anywayA null or negative result often reveals constraints you did not know you had. We tested a beautifully crafted explainer video on a SaaS pricing page. Engagement rose, time on page rose, but paid conversion did not budge. The post-test interviews clarified why. Prospects enjoyed the video but delayed the click to talk to sales until later. That told us two things. First, the video belonged upstream, in remarketing and nurture. Second, the pricing page is not the place for long attention work. The follow-up tests on the nurture path delivered a 9 percent lift in sales qualified leads at a lower cost per.
If your variant underperforms, examine variance across segments without p-hacking. You may find that new visitors respond poorly because the message assumes familiarity. That is a fixable scope issue, not a death sentence for the concept. Sometimes a losing test whispers, wrong audience, not wrong idea.
Running a portfolio without stepping on your own toesAs your program matures, coordination becomes the constraint. Parallel tests can interfere when they share traffic or when one changes the mix that the other depends on. Two homepage tests may appear independent, but if one shifts source mix toward mobile, the other’s effect changes. Keep a living map of concurrent tests, their randomization units, and the slices they touch. Traffic allocation tools help, but governance matters more. Stagger big bets. Bundle small tests that share a page area. Reserve shared components for dedicated windows.
Culture helps too. Reward teams for holding back when interference risk is high. Measure the throughput of valuable decisions per quarter, not the number of tests launched. A smaller portfolio with teeth is better than a wall of green boxes that move no revenue.
Telling the story so people act on itIf a result sits in a slide deck, it is dead. You have to publish it in the language your colleagues use to make decisions. A good readout starts with the decision question, shows the design briefly, presents the result in business units, then spells out the action with the rollout plan and monitoring. Put the data behind a link for the curious. Use visuals that show the distribution of outcomes, not just a single bar with a star.
Archive results in a way that will be searchable six months from now. Tag by channel, metric, and audience. It sounds bureaucratic, but it rescues teams from running the same test twice because the original project owner changed jobs. An org with institutional memory compounds learning. That is the essence of the uncommon part of (un)Common Logic. It is not a flourish, it is the quiet discipline to keep the knowledge flowing while people and platforms change.
Edge cases that separate rookies from prosA few patterns bite often enough that they deserve a final spotlight.
Promo cannibalization. Deep discounts lift conversion but often by shifting demand across time or from full-price channels. If your analytics cannot see halo and substitution across categories, do not trust basic per-visit revenue.
Auction dynamics. Creative that wins in a narrow A/B can lose in the wild because the auction mixes change. Re-run a subset of creative tests with budget caps mimicking production to test for scalability.
Learning decay. Some algorithmic systems adapt slowly. A test that toggles elements too quickly can produce results that vanish on rollout because the system never reached a steady state. Respect platform memory.
Identity drift. Cross-device users break cookie-level tests. If mobile web and app both contribute to conversion, align identification or move to geo or account-level randomization.
Delayed harms. A pricing test that lifts signups can backfire if it affects support burden or churn. Add delayed guardrails, even if you have to analyze them with a lagged cohort and a separate follow-up checkpoint.
The mindset behind the methodTools will change, privacy norms will evolve, platforms will tilt the board. The core of reliable marketing experiments does not change. Define what you are deciding. Randomize where signal is clean. Measure what matters, and protect the parts of the business that make the win sustainable. Size for effects that justify action. Commit to the rules before the heat of the moment. Explore with curiosity, confirm with restraint. Treat every test as a step in a longer conversation with your market, your systems, and your team.
That is what I mean by marketing experiment design with (un)Common Logic. It is not a slogan. It is the work of asking the annoying questions at the right time, so that your future self does not inherit a mess wrapped in a green arrow. When you hold to it, the wins come, and they stay won.