1 Foundational Ideas in Causal Inference
1.1 What’s New?
Deleted, Updated, and New Chapters
This is a new edition of a book I published in 2021. It may look like the first book, but it has been reasonably augmented and expanded upon. Some chapters have stayed largely the same, some have been updated significantly, and some are entirely new. Below, I’ll walk you through what’s changed and why.
Some chapters didn’t change much. The Directed Acrylic Graphs (DAG) chapter has not been radically changed, though I tried to more strategically connect it to the Unconfoundedness chapter than before. The Fixed Effects chapter is largely the same, though now it leads a new part of the book called Causal Panel Designs.
Other chapters, however, have undergone slight, but not radical, updates. The Instrumental Variables chapter has updated material on newer developments in diagnosing weak instruments, new examples, and other refinements. The Regression Discontinuity chapter features a new data exercise and additional updates to reflect modern practices in the field. The Potential Outcomes and Randomization chapter hasn’t changed dramatically but has been expanded with a new data exercise, slightly different framing, and some restructuring.
After that, things did change quite a bit.
I have removed the Regression and Probability chapter entirely. I relocated relevant portions of regression to the Unconfoundedness chapter, which has been heavily rewritten as well. This chapter now contrasts regression with matching and weighting techniques. I tried to emphasize simple things for thinking about regression and matching/weighting, such as trading off concepts like common support and functional form, extrapolation versus interpolation, and the role of heterogeneous treatment effects in clouding interpretation. I’ve worked to illustrate these ideas through simulations, code, and exercises, showing why traditional ordinary least squares (OLS) specifications might be biased but other specifications of OLS might not. To help readers navigate these complexities, I’ve included a checklist for applying unconfoundedness in empirical work. I have also removed all rap lyrics and rap examples from the book, though not by preference. I simply had a much harder time, for unknown reasons, getting the permissions.
The most significant expansions are in what I now call the “Causal Panel Designs” chapters. The difference-in-differences material has been completely overhauled. Instead of one chapter, there are now two:
The first chapter focuses on the fundamentals of difference-in-differences, walking readers step-by-step through a simple design with one treatment group. It emphasizes identification using potential outcomes and explains key assumptions like parallel trends and no anticipation. Through examples, I show why relying on an already-treated control group can be risky. I also provide a detailed list of evidence types that can support causal claims in this framework, including the event study and various falsification exercises.
The second chapter addresses complex designs, tackling issues like imbalanced panels, compositional changes, triple differences, and covariates. It also includes an in-depth discussion of differential timing and a ten-point checklist for applying difference-in-differences. I illustrate these concepts using a running example: evaluating the effects of Brazlian mental healthcare reform on murders.
The Synthetic Control chapter dives deeper into the biases introduced by imperfect fit, along with potential fixes like demeaning and augmented synthetic control. I’ve also introduced matrix completion with nuclear norm regularization—a more general panel estimator that accommodates a variety of designs, from single treatments to differential timing setups. To round it out, I’ve added a discussion of the increasingly popular synthetic difference-in-differences method.
Throughout the book, I have tried to emphasize two things: the importance of the treatment assignment mechanism and the importance of clearly defined target parameters expressed using potential outcomes, populations, and averages. I tried to stress practical things, too, like checklists and “no peeking.” And as before, I include coding exercises, and scattered throughout the book is code for Stata and R, with pointers to the online book at https://mixtape.scunning.com where you can find more complete code, as well as an updated errata page as errors are caught and fixed.
A Broader Vision for the Reader
I do not assume that readers are American. That is why I explain landmarks or references that American readers recognize but which others may not. I presume that not all readers come from economics departments or elite American universities, work in academia, or have extensive backgrounds in math and statistics. But I also try to tell the story as I understand it, and the story as I understand it is that this material, call it design-based causal inference, is the depositing of two great rivers into a much larger, much greater river. Those two smaller rivers are:
The empirical labor economics work of the Princeton Industrial Relations Section, particularly Orley Ashenfelter and his many academic children, grandchildren, and great-grandchildren such as the Nobel Laureates David Card and Josh Angrist, the late Alan Krueger and Bob Lalonde, the Nobel Laureate James Heckman, and the influential health, labor, and family economist Janet Currie
And the Harvard statistics tradition under Don Rubin and the Harvard economics program, particularly through the econometrician Guido Imbens
By emphasizing these origins, I hope readers see that there’s a sociological reason why some of this material is novel and strange. If you weren’t connected to either tradition, if you weren’t a labor economist, if you weren’t living in Boston or a certain section of the Princeton library in the 1970s and 1980s, there’s a decent chance that this material skipped right over you, and that by the time it did reach you, you felt like you were walking into the middle of a movie and being shushed when you asked who this person was or what the plot was even about. My hope is that these narratives and the historical context are not a distraction but rather help new readers take a step into this material with more confidence. I hope they provide a clearer map of this intellectual river, with all its meanders, rapids, invasive species, and oxbow lakes.
A Personal Note
Finally, I want to say this: I wrote this book for people who, like me, felt left out of these currents. The primary audience is the people who have at some time or another hesitated to raise their hand in econometrics classes, who are not economists nor do they want to be economists, who aren’t students in elite universities but who are dying to write the papers inside them. I love this material, and I think there are a lot more people like me out there who will too.
Thus the book is overwritten. It moves between pictures, stories, metaphors, and description of papers to technical statistics, simulations and code, common sense and intuition, and simple exercises. I think this is how a lot of people will learn this material. I think this is how to accommodate where people are now, rather than demanding they be somewhere they aren’t. Taking the long way around to learn this, like I did, is unnecessary and inefficient, and I hope this book shortens that path for you.
I hope this book reminds you that you belong—but of course, you do not need me, or this book, to remind you that you belong. Still, although this is a book about econometrics and statistics, that’s what I hope it does.
This book is not a textbook. This book is a mixtape. Mixtapes are interpersonal amateur creations given by one friend to another. They are a collection of favorite songs. These are some of my favorite songs for you.