3  Planning and designing credible causal analyses

Published

December 27, 2025

Modified

September 26, 2026

Keywords

A/B Testing, Causal Inference, Causal Time Series Analysis, Data Science, Difference-in-Differences, Directed Acyclic Graphs, Econometrics, Impact Evaluation, Instrumental Variables, Heterogeneous Treatment Effects, Potential Outcomes, Power Analysis, Sample Size Calculation, Python and R Programming, Randomized Experiments, Regression Discontinuity, Treatment Effects

Get news about the book. Leave your address here.
Found an error or have a suggestion? Submit feedback here.
Follow the book’s page for new content and follow the author.

Listen to a podcast version of this chapter, generated by Google’s NotebookLM.1

As we saw in Chapter 2, neither vast amounts of raw data nor advanced algorithms can answer a causal question on their own. Methods labeled Causal AI, Double ML, Causal Forest, or Meta-learners do not give you a free pass to claim causality. Data and methods are means to an end, and they cannot rescue the wrong question.

In this chapter, I’ll show you how to put theory into practice by working through a causal data project, from formulating a precise question to executing the analysis and interpreting the findings.

I’ll show you how to defend your conclusions against common sources of bias so your findings are statistically sound, sensible, and actionable. Throughout the chapter, we’ll again use simple examples from tech and digital businesses.

3.1 Formulating a well-defined causal question

“A problem well stated is a problem half solved.”
— John Dewey (North-American philosopher)

Every causal data project should begin with a precise question. Stakeholders often ask vague questions such as “Does X work?” or “What is the impact of Y?” These questions do not specify what, who, or how. Before you start, define the causal question as a testable statement about a cause-and-effect relationship.

A well-formed causal hypothesis has these components:

  1. What is the treatment/intervention in question? (\(D_i\)): What specific action, policy, or feature are you introducing or changing? This must be clearly defined and measurable. For instance, is it “sending a push notification” or “displaying a new UI element”?

    • What constitutes the intervention?: Define exactly what it means to receive the treatment. For example, is someone treated when they receive a discount email, whether or not they open it, or only when they open the email and click a link? This choice changes the effect you measure and how you interpret it (see Section 2.5). The right definition depends on the context, so make it explicit.
  2. What’s the KPI/outcome over which you want to measure an effect? (\(Y_i\)): What specific metric or behavior do you expect to change as a result of the treatment? This should be a key performance indicator (KPI) relevant to your business goals. Examples include “daily active users,” “average order value,” or “churn rate”. You may measure multiple KPIs, but specify which one is your focus.2

    • Be precise about outcome definitions: Different teams often define the same metric differently. For example, “user churn” might mean 30 days of inactivity for the marketing team but 90 days for the commerce team. Choose the definition that best aligns with your specific business question and stick to it consistently throughout your analysis.
  3. What’s your population and time period of interest? Who are the individuals, users, or entities that are subject to this treatment and whose outcomes you are measuring? Be specific. “All users” is rarely precise enough; remember there will be inactive users, for instance (e.g., those who haven’t accessed the platform in more than a year). Is it “new sign-ups in the last 30 days,” “high-value customers in the last year”, or “users in a specific geographic region during december 2024”?

  4. What is the comparison group (the “feasible” counterfactual, as discussed in Chapter 2): What is the alternative scenario against which you are comparing the treatment? This comparison is necessary for establishing causality. Is it “users who received the old feature”, “users who received no intervention,” or “users who received a placebo”?

Let’s use this framework to refine a common business question. The first version is how stakeholders may phrase it (“Vague question”); the second is the question you should help them reach (“Well-defined question”):

Vague question: “Does our new feature reduce churn?”

Well-defined question: “Among users active for at least 30 days before the experiment (population in question), does exposure to the new personalized feature (treatment), compared to the standard interface (comparison), reduce their likelihood of leaving the platform over the next 90 days (outcome)?”

Notice how this refined question immediately suggests what data you need to collect, who you need to target, and what you’ll compare. It also implicitly defines exactly what effect you want to measure (in this case, the average impact on new users).

3.1.1 Avoiding hypothesizing after the results are known

Once you have a well-defined causal question, pre-register the hypothesis and analysis plan: document the question, metrics, method, and decision rules before you collect or analyze the data. A slide deck used to pitch the project and validate it with stakeholders may be enough.

Pre-registration helps you avoid HARKing, or “Hypothesizing After the Results are Known”. HARKing means formulating a hypothesis after seeing the results, often subconsciously, to make findings appear more significant or fit a narrative. Even when unintentional, it can produce misleading conclusions and erode the credibility of your analyses.

Here’s a subtle example: An analyst is asked to “look into” whether a recent UI redesign affected user behavior. No specific hypothesis is written down, no primary metric is pre-registered. She dives into the data and finds… nothing dramatic. Conversion rates are flat. Bounce rates are unchanged. Session duration? Meh.

But then she notices something: users who saw the redesign scroll slightly farther down the page. “Aha!” she thinks. “The redesign increases content exploration!” She writes up a positive report, genuinely believing she found what she was looking for. Except she wasn’t looking for scroll depth when she started — it just happened to be the metric that moved. Without a pre-registered hypothesis, she had no way to distinguish between a real finding and a lucky pattern in the noise. This is HARKing as self-deception: an honest analyst fooling herself because she never committed to a specific question upfront.

Now consider a more blatant variant. A product team launches the XYZ algorithm to personalize their social feed. After running an A/B test, they discover mixed results — a modest 8% increase in total time spent in the app, but a concerning 15% drop in click-through rates for their most valuable user segments.

Under pressure to show success, the team changes its success metric. Suddenly, “engagement” no longer includes click-through rates — it’s redefined to mean only “time spent”. The report says, “Algorithm XYZ delivers breakthrough engagement gains!” while the CTR decline is buried in an appendix or omitted. This is HARKing as a cheating device: the team knew the results were mixed and hid the result that contradicted its preferred conclusion.

Both variants produce the same problem: false discoveries presented as real insights. The first may be harder to detect because the analyst did not intend to deceive anyone; she deceived herself. The second is cynical, but someone in the room knows the full picture. Pre-registering the question and analysis plan with the team and stakeholders before the experiment can prevent both.

Figure 3.1 compares HARKing to shooting arrows at a blank wall and painting bullseyes around wherever they land. The result is more false positives, resources spent chasing phantom wins, and a gradual loss of trust in your data practice. When stakeholders discover the full picture, rebuilding credibility can take years.

Figure 3.1: Hypothesizing after the results are known (HARKing) undermines the scientific integrity of causal analysis by allowing researchers to cherry-pick results and formulate hypotheses post-hoc, leading to false discoveries and unreliable conclusions.

Best Practice: Write down your causal question, estimand, and analysis plan in a shared document (e.g., a project brief) before the experiment or data collection begins. Doing so makes it harder to change the metric after seeing the results and keeps your conclusions tied to the pre-specified evidence.

3.2 Business therapy session: clarifying the project goals

Stakeholders rarely arrive with a well-defined causal question. They may ask for “causal insights” when they actually need a descriptive or predictive analysis. Before planning the project, help them clarify which kind of question they are asking. I call this the “business therapy session”.

Before planning the project, check whether the request sounds causal but isn’t. Consider this example:

Stakeholder request: “I want you to use our historical data to identify what characteristics make the sellers in our marketplace more prone to sales growth”.

Stakeholders often assume that studying high-performing sellers will reveal behaviors they can encourage others to copy. They are looking for an “aha moment”.

Strictly speaking, this request asks for a pattern in historical data, not a cause. Even rich historical data can only show associations, not causes.3

For instance, you might find that sellers with higher social media engagement have better sales growth. But this alone doesn’t tell you whether:

  • Social media engagement causes better sales (a causal relationship)
  • Better sales cause more social media engagement (reverse causality)4
  • Both are driven by a third factor like seller motivation or product quality (confounding)

Stakeholders want to know what behavior to encourage to achieve the desired outcome. Their question is: “If I incentivize sellers to engage more on social media, will their sales improve?” But to answer that, you would need an experiment or quasi-experiment (which you could then apply to historical data) that manipulated these characteristics and observed the results. Only then can you disentangle correlation from causation.

3.2.1 Four essential clarifying questions

I recommend sitting down with your stakeholders and partners to answer these four questions and those in Section 3.3. Getting the answers early helps keep the project on track and focused on a decision that matters.

To transform vague stakeholder requests into actionable causal projects, work through these four questions together:

1. What is the endgame of the project? Define “what success looks like” in objective terms. Most teams already have target metrics such as revenue, retention, or conversion rates. Ask which outcome the project should change and how the team will recognize success. This moves the conversation from “let’s explore the data” to a concrete, measurable goal.

2. What is the core question or hypothesis we are addressing? This builds on the endgame to formulate the primary hypothesis or question the project intends to answer. Is it clear, testable, and relevant to stakeholders and to the business? More importantly, is it actually a causal question (one asking whether X causes Y) that requires causal inference methods, or is it a descriptive/predictive question (summarizing patterns or forecasting outcomes) that can be answered with standard analytics?

3. What decisions or actions will be taken based on our insights? Identify what stakeholders will do if the result is positive, negative, or inconclusive. If no decision will change based on the analysis, question whether the project is worth pursuing.

4. How will we measure project success on our end? Define the metrics or KPIs that will show when your work is done. This sets boundaries for the analysis and prevents endless “deep dives” when stakeholders are unsure how to proceed. It also helps prevent HARKing (revising your hypothesis after seeing the results) by establishing success criteria upfront.

These questions serve as a filter: they help distinguish genuine causal questions from correlation/prediction questions, ensure stakeholder buy-in for acting on results, and establish clear boundaries for your analysis scope.

3.3 Gather more context to design effective causal studies

Once your question is locked in, design a study that can credibly answer it. The design must defend against the biases discussed in Chapter 2, particularly omitted variable bias and selection bias. The central challenge is to construct a valid counterfactual: a group that represents what would have happened to the treated group without the treatment.

Before choosing a method, ask which study designs are feasible and which biases each design must address. The same questions apply whether you’re planning a new experiment (we call it “ex-ante”) or analyzing an intervention that has already occurred (we call it “ex-post”).5

The discussion below previews several ways to design a credible causal study. I cover each one in depth later in the book. If you want an early summary of how the methods work, what they require, and how they support causal identification, see Appendix 3.A.

Has anyone received the treatment yet?

The answer determines whether you are planning a controlled experiment before exposure (ex-ante) or analyzing data from an intervention that already happened (ex-post). It also narrows the methods you can use and the biases you need to address.

If the answer is “no one has been exposed yet”, you may be able to run a Randomized Controlled Trial (RCT), including the familiar A/B test. Users are randomly assigned to a treatment group (exposed to the intervention, e.g., a new feature) or a control group (not exposed). Like shuffling a deck of cards, randomization makes the groups similar on average across both measured and unmeasured characteristics. This eliminates selection bias and omitted variable bias.

If the answer is “yes, some have been exposed at different periods”, this scenario often arises from partial rollouts, phased launches, or natural adoption of a feature. Unlike an RCT, it lacks random assignment, but it opens the door for quasi-experimental designs, methods that mimic experiments using observational data. Methods designed to reconstruct a credible comparison group become highly relevant, including difference-in-differences (Chapter 9, Chapter 10).

If the answer is “yes, everyone has been exposed at once”, causal inference becomes more difficult because there is no unexposed control group. In such cases, you often need to rely on designs with stronger assumptions, meaning more things have to be true (and harder to verify) for the analysis to be valid. For example, if a major social media platform rolls out a new feed algorithm globally, you might look for a natural discontinuity (a threshold or cutoff that creates a before/after comparison), use synthetic control methods (building a “virtual” control group from other data), or construct a counterfactual based on time series (Chapter 11).

Who decides who gets the treatment?

Your control over treatment assignment directly affects how credible the study can be. More control can bring the design closer to a randomized experiment, which offers the strongest causal evidence. Less control requires more complex methods to address potential biases.

If your answer is “yes, I can choose individually”, you can run a randomized experiment. Control over assignment is the strongest defense against selection bias because it makes the intervention the only systematic difference between groups. We examine this design in Chapter 4.

While individual-level randomization is often preferred, it’s not always practical or possible, especially for interventions that inherently affect groups (e.g., a new policy, a network-level change). So if your answer is “yes, I have control over who receives the intervention by groups (e.g., by region, store, time period)”, then cluster randomization (randomizing at the group level) or designs like difference-in-differences (comparing changes over time between groups) become relevant. For instance, a food delivery app might test a new checkout flow in certain cities, affecting all individuals within each city, and compare order completion rates to similar cities still using the old checkout process.

You have “partial control” when a rule, policy, or natural process influences treatment assignment in a way you can exploit. This setting may support instrumental variables (IV, Chapter 7) or Regression Discontinuity Designs (RDD, Chapter 8).

Think of an instrumental variable (IV) as a “nudge” that pushes people toward treatment without directly affecting the outcome in any other way. Formally, an IV influences treatment but only affects the outcome through that treatment, and it isn’t related to the hidden factors you can’t measure. For example, consider a company wanting to estimate the causal effect of attending a training program on employee performance. A random encouragement to attend the training (e.g., a randomly sent email invitation) could serve as an instrumental variable. The encouragement influences attendance but only affects performance through attendance, and is not directly related to unobserved factors affecting both attendance and performance.

RDD is applicable when treatment assignment is determined by a sharp cutoff point on a continuous variable. For example, a digital advertising platform might offer a premium feature only to advertisers who spent above a certain threshold in the previous month. By comparing advertisers just above and just below this threshold, you can estimate the causal effect of the premium feature, with the only systematic difference between groups being whether they received the treatment. This design works almost like a randomized experiment near the cutoff, because advertisers just above and just below the threshold are essentially similar, just on different sides of the line. That similarity provides strong causal evidence.

If you have “no control” at all over who receives the intervention, you’ll need to work with the data you have and look for naturally occurring differences in who received the intervention and when. One approach is to explore time variation in treatment: perhaps the intervention was rolled out at different times to different groups, or there were natural periods when some individuals were exposed while others weren’t. You can also account for observable user characteristics that may influence both treatment assignment and outcomes. Regression adjustment or matching can “balance” the groups on those characteristics, while difference-in-differences tracks how the groups change over time.

Could something else explain or influence your results?

External factors can affect the outcome and confound the estimate. Identify them before the study so you can decide how the design will handle them. The examples below are common competing explanations for your results:

Seasonality: Many digital businesses experience seasonal fluctuations in user behavior, demand, or engagement. For example, an e-commerce platform might see increased sales during holiday seasons, or a travel booking app might observe higher usage during summer months. If your intervention coincides with a seasonal peak or trough, attributing observed changes solely to your intervention would be misleading. To defend against this, you can:

  • Run your experiment during periods of stable seasonality or compare periods with similar seasonal patterns.

  • In observational studies, include time-based variables (e.g., month, quarter, day of week) as control variables in your regression models.

  • Use difference-in-differences, which by construction accounts for common time trends, including seasonality, as long as the seasonal patterns are similar between your treatment and control groups.

Concurrent marketing campaigns or product launches: Your company or competitors might be running other initiatives simultaneously with your study. A new marketing campaign could drive traffic to your website, or a competitor’s new feature could divert users away. These concurrent events can confound your results. To mitigate this:

  • If possible, align your study timing to avoid overlap with other major internal initiatives.

  • Keep track of competitor activities, industry news, and major market shifts that could impact your study. If such events occur, consider them as potential confounders and, if data is available, include them as control variables or analyze their impact separately.

  • If a concurrent campaign targets a specific segment of your user base, you might be able to exclude that segment from your study or analyze its impact separately.

Any unaccounted factor that affects both treatment assignment and the outcome can cause omitted variable bias. Identify these factors before the study and, where possible, address them through study design, data collection, or statistical modeling.

Can you establish a baseline and track changes over time?

Having data from before the intervention is one of your most valuable assets. It lets you establish a starting point, check whether your comparison groups were similar to begin with, and control for factors that could otherwise bias your results.

Establishing baselines: Pre-intervention data lets you measure the change attributable to your treatment, rather than just comparing absolute levels between groups, as shown in Figure 3.2. For example, if you’re testing a new feature to increase user engagement, knowing the baseline engagement for both treatment and control groups allows you to compare the change in engagement, not just the final levels. Without this baseline, observed differences might reflect pre-existing disparities rather than your intervention’s effect.

Figure 3.2: Establishing baselines with pre-intervention data

Checking that groups were on similar trajectories: Difference-in-Differences (DiD) designs rely on a key assumption: before the intervention, both groups were trending in a similar way. Economists call this the “parallel trends assumption” (also illustrated in Figure 3.2). Pre-intervention data lets you check this by plotting outcome trends for both groups before anything happened. If the trends were already diverging, you can’t trust your DiD results, because the groups weren’t comparable to begin with. For instance, if a fintech company tests a budgeting tool’s impact on savings by comparing treated and control regions, but savings trends were already different before the tool launched, the results won’t be credible.

Controlling for prior behavior: In observational studies, pre-intervention data helps control for selection bias by accounting for baseline differences between groups. You can include past behavior in your statistical models to account for baseline differences, or use it to find “look-alike” users for comparison. For example, when evaluating a new recommendation system, including users’ historical viewing patterns helps ensure you’re comparing similar users who did and didn’t adopt the new system.

Without pre-intervention data, you lose the ability to distinguish between changes caused by your intervention and pre-existing differences between groups, weakening your causal claims.

How long should you wait before measuring the effect?

Choose the observation period based on how quickly the intervention can affect the outcome. Measuring too early can miss a delayed effect, while measuring too late can waste resources and expose the study to more outside events. The expected timing should determine both the study duration and the metrics you track.

Immediate vs. lagged effects: Some interventions yield immediate results. For example, changing the color of a ‘Buy Now’ button on an e-commerce site might have an almost instantaneous impact on click-through rates or conversion. Other interventions, however, might have lagged effects. A new user onboarding flow designed to improve long-term retention might not show its full impact for weeks or months, because you need to wait and see whether users actually stick around. Similarly, the impact of a new content recommendation algorithm on user loyalty might only become apparent after users have interacted with the system for a sustained period.

Defining the observation window: Based on your understanding of the intervention and the expected user behavior, you need to define an appropriate observation window for your study. This involves:

  • Short-term metrics: For immediate effects, you might focus on metrics like click-through rates, conversion rates, or session duration. These can be observed within hours or days.

  • Long-term metrics: For lagged effects, you need to track metrics like retention, lifetime value (LTV), churn rate, or subscription upgrades over weeks, months, or even quarters. Running an A/B test for only a few days when the true impact is on long-term retention would lead to an incomplete and potentially misleading conclusion.

Getting the study duration right: Run your study too short, and you risk missing the real effect. Run it too long, and you waste resources, delay decisions, and give external factors more time to muddy the waters.

For example, a social media platform testing notifications for dormant users should run the experiment long enough to distinguish a brief return from sustained engagement.

Accounting for novelty effects: New features or designs can create a temporary ‘novelty effect’: users engage more because the experience is new, not because the change improves the product. This response typically fades. A long enough study helps distinguish it from a sustained causal effect.

For example, a new UI design might initially boost engagement, but if the effect diminishes after a few weeks, it might indicate a novelty effect rather than a fundamental improvement. In this case, a study that only runs for the first week after launch might conclude that the new design is highly effective, when in reality the improvement is temporary.

The takeaway: effects don’t always show up right away. As illustrated in Figure 3.3, the distinct shape of the effect over time matters just as much as its magnitude. You might encounter three very different trajectories:

Figure 3.3: The shape of the effect over time matters as much as its magnitude
  • Panel A: The short-lived effect. As the novelty wears off, the treatment group regresses back to the baseline trend of the control group. This is the classic “novelty effect” trap: if you measure too early, you see a lift; if you measure later, it’s gone.
  • Panel B: The sustained effect. This is the ideal scenario for a product launch. The intervention causes a divergence between the treatment and control groups that persists over time, confirming that you created lasting value.
  • Panel C: The delayed effect (“The slow burn”). After the intervention, there is a “dead zone” where the treatment and control lines remain identical. This often happens in complex products where there is a learning curve, or in network-effect features where a critical mass of users is needed before value is realized.

Knowing when to measure is as important as knowing what to measure. Get the timing wrong, and your conclusions may be incomplete or wrong.

3.4 Executing the analysis using the steps above

Imagine you’re a data scientist at a large ecommerce marketplace. Your product team is testing a new AI-based homepage layout, designed to personalize the first screen with machine-curated product recommendations. The hypothesis is straightforward: more relevant content should lead to more purchases.

To test that, you randomly assign 50% of users to see the new layout, while the other 50% keep seeing the classic homepage. The rollout is server-side, so there’s perfect compliance — no opt-ins, no app store versions to track. Everyone sees what they were assigned to. Your mission: estimate the causal effect of this change on user spending.

3.4.1 Defining the ingredients of the analysis

The treatment is exposure to the new AI-curated homepage: the user opens the app and sees a dynamically generated layout showing products the system believes they’re most likely to buy.

Your KPI is profit per user. But you need to be more precise. So here’s how you define it:

  • It’s the user-level profit (revenue minus variable costs) generated from all completed purchases, in R$ (Brazilian real), excluding refunds and canceled orders.

  • It only includes purchases initiated after the user saw the homepage, because our hypothesis is that the new layout is the mechanism or channel that leads to more purchases.

  • It’s measured from day 0 (the exposure date) to day 30, to capture both quick and delayed effects. We could even extend the window later to check whether the effect persists over time or if it’s just a novelty effect.

  • Our population: we’re focusing on users in Brazil who opened the app during the experiment window.

This randomized experiment with perfect compliance has a clear feasible counterfactual: users assigned to the classic homepage. Because randomization balances other factors on average, differences in outcomes between the groups can be interpreted causally.

3.4.2 Lining up the business context

Before opening your favorite R or Python IDE, you recall what I discussed in Section 3.2. You called your main stakeholders and asked them: What’s the endgame here?

The goal is to decide whether to roll out the new homepage layout to all users in Brazil. If the answer is yes, stakeholders also need the expected additional profit per user and per month to calculate the return.

Stakeholders specify: “If the uplift is statistically significant and at least R$0.8 per user, then we’ll push the new layout to everyone.” The R$0.80 threshold is their minimum effect of interest (MEI): the smallest uplift that would justify the cost of building and maintaining the personalization engine and produce a positive Return on Investment (don’t worry about how to calculate ROI yet. I’ll cover that in Chapter 14). If the estimate is weak or statistically inconclusive, the design team will reconsider the layout.

Most users who convert after visiting the homepage do so within 48 hours, but the plan still measures 30-day cumulative profit. This window protects against early spikes or novelty bias and captures both fast and slow shoppers.

If the effect is positive, revisit the analysis later to test whether it persists over longer time horizons.

3.4.3 Getting to the numbers

Data: All data we’ll use are available in the repository. You can download it there or load it directly by passing the raw URL to read.csv() in R or pd.read_csv() in Python.6

Packages: To run the code yourself, you’ll need a few tools. The block below loads them for you. Quick tip: you only need to install these packages once, but you must load them every time you start a new session in R or Python.

# If you haven't already, run this once to install the package
# install.packages("tidyverse")
# install.packages("RCT")

# You must run the lines below at the start of every new R session.
library(tidyverse)
library(RCT) # Implements the RCT tools
# If you haven't already, run this in your terminal to install the packages:
# pip install pandas numpy statsmodels scipy.stats (or use "pip3")

# You must run the lines below at the start of every new Python session.
import pandas as pd # Data manipulation
import numpy as np # Mathematical computing in Python
import statsmodels.formula.api as smf # Linear regression
from scipy.stats import ttest_ind

Start by connecting the business question to the available data. Inspect the variables and their definitions:

  • treatment: Binary treatment indicator (0/1) for whether user received the personalized feed treatment
  • profit_30d (outcome): Total profit generated by each user in the 30 days following the intervention (R$)
  • pre_profit: Total profit generated by each user in the 30 days prior to the intervention (R$)
  • user_tenure: Number of months each user has been active on the platform
  • prior_purchases: Total number of purchases each user made on the platform before the intervention

Because assignment is randomized and compliance is perfect (i.e., users assigned to the personalized feed inevitably receive it when they visit the homepage), estimate the Average Treatment Effect (ATE) by comparing mean user-level profit between groups. The code below first checks covariate balance, then estimates the effect in R and Python.

# read the CSV file
df <- read.csv("data/personalized_feed_experiment.csv")

# balance table between treatment and control groups with t-tests for all variables
balance_table <- balance_table(
  data = df,
  treatment = "treatment"
)
print(balance_table)

# main result
main <- lm(profit_30d ~ treatment, data = df)
summary(main)
confint(main) # get confidence intervals
# read the CSV file
df = pd.read_csv('data/personalized_feed_experiment.csv')

treatment_col = 'treatment'
g1 = df[df[treatment_col] == 0]
g2 = df[df[treatment_col] == 1]

# balance table between treatment and control groups with t-tests for all variables
balance_table = pd.DataFrame([
    {
        'variable': col,
        'mean_control': g1[col].dropna().mean(),
        'mean_treat': g2[col].dropna().mean(),
        't-stat': ttest_ind(g1[col].dropna(), g2[col].dropna()).statistic,
        'p-value': ttest_ind(g1[col].dropna(), g2[col].dropna()).pvalue
    }
    for col in df.select_dtypes(include='number').columns if col != treatment_col
])
print(balance_table)

main = smf.ols('profit_30d ~ treatment', data=df).fit()
print(main.summary())

Baseline comparison and covariate balance: The balance table generated by the code tab above and shown in Table 3.1 compares pre-treatment characteristics — like user_tenure, prior_purchases, and pre_profit — between groups. As expected in a well-executed randomized experiment, none of the imbalances is statistically significant (all p-values are greater than 0.05). This supports the claim that post-treatment differences are likely caused by the intervention.

Table 3.1: Balance table checking for pre-treatment differences
Variable Control Mean Treatment Mean p-value
pre_profit 22.6 22.5 0.843
prior_purchases 4.98 5.01 0.714
user_tenure 18.8 18.3 0.091

Estimating the treatment effect: Our first model, in that same codeblock, is a simple linear regression (profit_30d ~ treatment). This gives us a point estimate of R$2.34: users who saw the new homepage generated, on average, R$2.34 more profit over 30 days than those in the control group. The 95% confidence interval ranges from R$1.75 to R$2.93, and the p-value is near zero — indicating strong statistical significance.

Robustness checks: In Chapter 13, I’ll show you how to stress-test results like these. For now, run two robustness checks for whether confounding factors are driving the estimate. The code tab below implements them.

# controlling for user history: features do not account for the treatment effect
robustness1 <- lm(profit_30d ~ treatment + user_tenure + prior_purchases + pre_profit, data = df)
summary(robustness1)

# pre-experiment profit: no effect before the treatment
robustness2 <- lm(pre_profit ~ treatment, data = df)
summary(robustness2)
# controlling for user history: features do not account for the treatment effect
robustness1 = smf.ols(
    'profit_30d ~ treatment + user_tenure + prior_purchases + pre_profit',
    data=df
).fit()
print(robustness1.summary())

# pre-experiment profit: no effect before the treatment
robustness2 = smf.ols('pre_profit ~ treatment', data=df).fit()
print(robustness2.summary())

The first check, model robustness1, verifies that our treatment effect remains stable when we control for baseline characteristics. The effect size holds steady: R$2.44, with similar precision. This suggests that the uplift isn’t being driven by imbalances in user history. If the effect had disappeared or changed dramatically, it might suggest our randomization wasn’t perfect or there are confounding factors.

The second check, model robustness2, ensures there was “no effect before the treatment began”. Look what you did there: you regressed pre-treatment profit on treatment assignment and found no effect (point estimate = -R$0.07, p-value = 0.84). That is, the treatment has no impact on past behavior (as it shouldn’t), reinforcing the causal interpretation. If we had found a significant “effect” on pre-treatment outcomes, it would be a red flag indicating potential randomization failure or selection bias.

Takeaway to stakeholders: The new AI-curated homepage drove an average increase of R$2.34 per user in 30-day profit (95% CI: R$1.75 to R$2.93). The effect is statistically significant, holds after controlling for user history, and passes placebo tests. If rolled out to all users in Brazil, this translates into a projected profit uplift of R$11.7 million/month, assuming a 5 million monthly active user base.

3.5 Wrapping up and next steps

We’ve covered a lot of ground in this chapter. By now you can:

  • Formulate a well-defined causal question: move from vague stakeholder asks to precise, testable hypotheses that define the treatment, outcome, population, and comparison group.

  • Clarify the business goals before diving into data: you’ve learned how to align on what decisions your analysis will inform and what success looks like.

  • Tell whether you will plan an RCT (ex-ante) or analyze past data (ex-post), and what methods are feasible for each case.

  • Anticipate and defend against key biases: You’ve seen how omitted variables, selection bias, seasonality, and external shocks can threaten your analysis.

  • Execute an analysis of experimental data: estimate the effect of a personalized homepage in an A/B test with perfect compliance, then check the result with robustness and placebo tests. For now, we have focused on this ideal setting to build intuition.

This chapter completes the first part of the book — dense enough that we could have called it “Knowing this is half the battle”. The next part shows how to estimate impacts and report results as we did in Section 3.4.3, but with methods that also cover settings without a randomized experiment. If this is the last chapter you read, Appendix 3.A previews those methods.

Appendix 3.A: A preview of the methods we will cover

Randomized experiments (such as A/B tests)

In tech, the most common and powerful design for establishing causality is the Randomized Controlled Trial (RCT), often called an A/B test. Randomization balances observed and unobserved characteristics between treatment and control groups on average. As a result, outcome differences can be attributed to the treatment rather than pre-existing differences.

How it works:

  1. Random assignment: Users (or other units of analysis) are randomly assigned to either the treatment group (A) or the control group (B). This is the critical step that creates comparable groups.
  2. Treatment application: Group A receives the new feature/intervention, while Group B receives the old version or no intervention.
  3. Outcome measurement: You measure the specified outcome metric for both groups over a defined period.
  4. Comparison: You compare the average outcome between Group A and Group B. The difference is your estimated causal effect.

Example: A/B Testing a New Search Algorithm

Let’s say a large e-commerce platform wants to test if a new search algorithm improves conversion rates (users making a purchase after searching).

  • Causal Question: Among users performing a search, does exposure to the new ML-powered search algorithm (treatment), compared to the existing keyword-based algorithm (control), increase the probability of making a purchase within the same session (outcome)?
  • Population: All users performing a search on the platform.
  • Design: A/B test. Randomly assign 50% of search queries to the new algorithm (Treatment) and 50% to the old algorithm (Control).

Why it defends against bias: Random assignment breaks the link between treatment and potential confounders such as tech-savviness, income, or prior purchase history. On average, users in the new algorithm group are no more likely to convert before treatment than users in the old algorithm group. This removes selection bias and omitted variable bias.

Observational studies and quasi-experiments

While A/B tests are ideal, they aren’t always feasible due to ethical, technical, or business constraints. In such cases, you’ll rely on observational studies or quasi-experiments. These designs attempt to mimic the conditions of an A/B test by carefully controlling for confounding variables or exploiting naturally occurring randomization. Here are some common quasi-experimental designs and how they help address bias.

Difference-in-Differences (DiD)

Intuition: Imagine you want to measure the impact of a new pricing strategy launched in one city, but not in others. A simple comparison of sales before and after the change in the treated city might be misleading due to overall market trends (e.g., seasonality). DiD addresses this by comparing the change in the outcome in the treated group to the change in the outcome in a control group that did not receive the treatment.

How it defends against bias: DiD controls for unobserved factors that change over time and affect both groups equally (e.g., macroeconomic trends, overall market growth). It assumes that, in the absence of the treatment, the trend in the outcome would have been the same for both the treated and control groups (the ‘parallel trends’ assumption).

Example: Impact of a subscription price increase

A streaming service increases its monthly subscription price in Region A but keeps the price the same in Region B. We want to estimate the causal effect of this price increase on subscriber churn.

  • Treatment: Price increase in Region A.
  • Outcome: Monthly churn rate.
  • Population: Subscribers in Region A (treated) and Region B (control).
  • Design: Measure churn rates in both regions for a period before the price change and a period after the price change.

Regression discontinuity design (RDD)

Intuition: RDD is used when treatment assignment is based on a sharp cutoff point of a continuous variable. For example, if users with a credit score above 700 get a special loan offer, while those below 700 don’t. The idea is that users just above and just below the cutoff are very similar in all unobserved characteristics, except for their treatment status.

How it defends against bias: By focusing on units very close to the cutoff, RDD effectively creates a quasi-randomized experiment. Any differences in outcomes between those just above and just below the threshold can be attributed to the treatment, as other confounding factors are assumed to vary smoothly around the cutoff.

Example: Impact of a premium feature unlock

A SaaS company offers a premium feature for free to users who have spent more than 10 hours in their app during the first month. We want to estimate the causal effect of accessing this premium feature on user engagement (e.g., daily active usage).

  • Treatment: Access to premium feature.
  • Outcome: Average daily active usage (DAU) in the second month.
  • Population: Users whose first-month usage is close to the 10-hour cutoff.
  • Design: Compare DAU for users just above 10 hours to those just below 10 hours.

Instrumental variables (IV)

Intuition: IV is used when you suspect that your treatment is influenced by hidden factors that also affect the outcome. An instrumental variable is a third variable that influences the treatment but does not directly affect the outcome, except through its effect on the treatment. It acts as a ‘natural experiment’ that randomly assigns treatment.

How it defends against bias: The IV method isolates the “good” variation in the treatment (the part caused only by the experiment). This isolated variation is then used to estimate the causal effect for the group influenced by the instrument, effectively removing the bias caused by unobserved confounders.

Example: Impact of ad exposure on purchase conversion

Suppose we want to measure the causal effect of seeing an ad on purchase conversion. We know that users who see ads might be inherently more engaged or more likely to purchase (selection bias). However, a technical glitch causes some users to randomly not see ads, even if they were supposed to. This glitch can serve as an instrumental variable.

  • Treatment: Ad exposure (binary: 1 if user saw ad, 0 otherwise).
  • Outcome: Purchase conversion (binary: 1 if user purchased, 0 otherwise).
  • Instrumental Variable: Technical glitch (binary: 1 if glitch prevented ad, 0 otherwise). The glitch affects ad exposure but should not directly affect purchase conversion other than through ad exposure.
  • Population: All users.

Time series approaches

Intuition: When you don’t have a control group because everyone receives the intervention at the same time (e.g., a nationwide policy change or a global feature rollout), you can use the pre-intervention history to predict what would have happened if the intervention hadn’t occurred. The “counterfactual” is essentially an extrapolation of the past into the future.

How it defends against bias: By modeling the pre-existing trend and seasonality, you can subtract these factors from the post-intervention outcome. This allows you to distinguish the effect of the intervention from the natural evolution of the metric, assuming that no other major events occurred at the exact same time.

Example: Impact of a TV ad campaign

A company launches a nationwide TV ad campaign and wants to know if it drove an increase in website visits. Since the ad aired everywhere simultaneously, there is no unexposed control group.

  • Treatment: Launch of the TV ad campaign.
  • Outcome: Daily website sessions.
  • Population: All potential visitors (nationwide).
  • Design: Use an Interrupted Time Series (ITS) or CausalImpact model to predict the expected number of sessions based on historical trends (counterfactual) and compare it to the actual observed sessions.

Choosing the right design

The choice of causal design depends heavily on the context, data availability, and the nature of the intervention. Here's a quick guide:

  • A/B Test: Always prefer if feasible. It's the most robust against unobserved confounders.
  • Instrumental Variables: Useful when treatment is influenced by user choice or hidden factors, and you can find a valid instrument that affects treatment but not outcome directly.
  • Regression Discontinuity: Applicable when treatment assignment is based on a sharp cutoff of a running metric (like a credit score, age, or total spend).
  • Difference-in-Differences: Good when you have panel data (before/after, treatment/control groups) and can assume parallel trends.
  • Time series approaches: Useful when everyone receives the treatment at the same time (no control group), but you have a long history of data before the intervention to model the counterfactual.

Regardless of the method, the goal is to construct a comparison that approximates the counterfactual: what would have happened to the treated units without treatment. Planning that comparison before the analysis is your strongest defense against a misleading conclusion.


  1. As Google advises, “NotebookLM can be inaccurate; please double-check its content”. I recommend reading the chapter first, then listening to the audio to reinforce what you’ve learned.↩︎

  2. There is an issue about testing multiple KPIs at once: the more KPIs you test, the more likely you are to find a spurious effect. Therefore, whenever you are testing multiple KPIs, you should use a multiple testing correction method.↩︎

  3. If the stakeholder instead wanted to predict which sellers will grow fastest, a predictive model would work fine. That’s a different project entirely.↩︎

  4. You will find more about reverse causality in the Appendix 2.A of the previous chapter.↩︎

  5. Ex-ante (forward-looking) means you’re designing the study before the intervention happens; ex-post (backward-looking) means you’re analyzing data from an intervention that already occurred. Ex-ante designs generally produce more credible causal evidence because you can plan for proper causal identification upfront.↩︎

  6. To load it directly, use the example URL https://raw.githubusercontent.com/RobsonTigre/everyday-ci/main/data/advertising_data.csv and replace the filename advertising_data.csv with the one you need.↩︎