Recent enthusiasm for field experiments, and especially for natural field experiments (NFEs), in which subjects go about their daily activities unaware that any study is taking place, has sometimes been read as a verdict against the laboratory. I argue that such a verdict is wrong. Through the lens of a simple rational-choice model, I show that the four standard experimental designs (laboratory, artefactual field experiment (AFE), framed field experiment (FFE), and NFE) are comparative-static restrictions of one maximization problem, each identifying a parameter the others cannot. The model reveals that each design type has distinctive strengths and weaknesses across various dimensions of knowledge creation, including the enforcement of the conditions for causal identification, the faithfulness of the experimental environment to the theory being tested, the identification of economic primitives via theoretical structure, and the ethics of studying human subjects. On each of the dimensions, the four design types are complements rather than rivals. Nowhere is this complementarity more evident than between the two extremes. The lab enforces the conditions for causal identification that the NFE must inherit from the market; the NFE recovers the parameter that governs behaviour in the wild, free of the selection, scrutiny, and environmental distortions the lab cannot escape. A research programme using all four designs together demonstrates something no single design can produce. The framework further accommodates the recent rise of online and survey experiments as natural extensions. Our discipline’s recent drift away from laboratory evidence is leaving an important structural gap that natural field experiments, however well conceived, cannot fill.
Aaron Bodoh-Creed, Brent Hickman, John List, Ian Muir, and Gregory Sun
Nonlinear pricing theory predicts that firms can extract surplus by inducing heterogeneous consumers to self-sort across price contract offers that are ex-post optimal for them. We study subscription pricing when the frictionless sorting assumption fails. Using large-scale subscription experiments conducted by Lyft, we document systematic deviations from optimal self-selection: many high-demand consumers decline subscriptions that would have saved them money, while some subscribers fail to break even. We develop a structural model of intensive-margin demand in which consumers may exhibit salience failures, forecast errors about future demand, or impulsivity. We show that subscription uptake can be recast as one-sided noncompliance in a binary-instrument framework, allowing us to leverage LATE methods to identify counterfactual outcome distributions and a novel "uptake function" linking baseline outcomes to compliance behavior. Combining experimental price variation with this identification strategy, we recover utility primitives, demand heterogeneity, and behavioral parameters. Salience failures and forecast errors play quantitatively important roles. Counterfactual analyses show that optimal subscription pricing generates substantial gains relative to linear pricing, but these gains are highly sensitive to consumer deviations from ex-post optimal choice. Implementing nonlinear pricing therefore requires not only optimal contract design for consumer screening, but also coordinated efforts to mitigate behavioral frictions.
In 2019 I put together a summary of data from my field experiments website that pertained to framed field experiments (see List 2024; 2026). Several people have asked me if I have an update. In this document I update all figures and numbers to show the details for 2025. I also include the description from the 2019 paper below.
In 2019, I put together a summary of data from my field experiments website that pertained to natural field experiments (Harrison and List, 2004). Several people have asked me for updates. In this document I update all figures and numbers to show the details for 2025. I also include the description from the original paper below.
List Experiments are widely used across the social sciences to measure sensitive attitudes and behaviors, yet no prior study has validated their estimates against an incentive-compatible behavioral measure. I conduct a field experiment with 400 subjects at a sports card show, combining List Experiment treatments for willingness to pay, one for wolf reintroduction in Yellowstone Park, one for a graded sports card, with a Vickrey second-price auction that provides a real-money benchmark. The List Experiment estimates 26% would pay $50 for the card, compared to 22% who bid at least that amount in the auction; this difference is not statistically significant. These results provide the first criterion validity test of a List Experiment and suggest the method holds promise as a parsimonious alternative to conventional stated preference approaches in settings where survey space constraints preclude standard bias-mitigation interventions.
Many ideas show remarkable returns in small-scale trials but often disappoint when scaled to broader populations and contexts. Using early childhood investment as a case study, this study develops a dynamic human capital formation model that integrates complementary skill investment with "Option C thinking" on scaling challenges. The model is stylized in the Chicago tradition: micro-founded with optimizing agents, dynamic skill production, and a policymaker evaluating scaling decisions. It formalizes how naive extrapolation from pilot studies systematically overestimates policy efficacy by ignoring "voltage drops," declining treatment effects due to unrepresentativeness at scale. The model demonstrates that optimal scaling policy requires mechanism-based design that anticipates these failures through backward induction from implementation realities. The scientific insights from a set of recent studies provide valuable perspectives on the model.
In 2019, I put together a summary of data from my field experiments website that pertained to artefactual field experiments. Several people have asked me if I have an update. In this document I update all figures and numbers to show the details for the year 2025. I also include the description from the 2019 paper below. The definition of artefactual field experiments comes originally from Harrison and List (2004) and is advanced in List (2006; 2024; 2026).
Concealing candidate identities during evaluations ("blinding") is often proposed to combat discrimination, yet its effects on the composition and quality of selected candidates, as well as its underlying mechanisms, remain unclear. I conduct a field experiment at an international academic conference, randomly assigning all 657 submitted papers to two blind and two non-blind reviewers (245 total) and collecting paper quality measures---citations and publication statuses five years later. I find that blinding significantly shrinks gaps in reviewer scores and acceptances by student status and institution rank, with no significant effects by gender. These increases in representation are not at the expense of quality: papers selected under blind review are of comparable quality to those selected non-blind. To understand mechanisms, I run a second field experiment that again implements blind and non-blind review, and elicits reviewer predictions of future submission outcomes. I combine my experiments to estimate a model of reviewer scores that uses blind scores to decompose non-blind disparities into distinct forms of discrimination. I find that the nature of discrimination differs by trait: student score gaps are explained by inaccurate beliefs about paper quality (inaccurate statistical discrimination) and alternative objectives (such as favoring authors whose acceptance benefits others), while institution gaps are attributable to residual drivers of discrimination such as animus.
Caroline Gaudreau, Dani Levine, John List, and Dana Suskind
Research shows responsive caregiving enhances children's brain development, with parental knowledge predicting positive behaviors and outcomes. However, knowledge varies widely across educational levels, highlighting the need for targeted interventions. Despite evidence that this knowledge can be improved, no comprehensive metric exists for efficient assessment. We introduce SPEAK (Survey of Parent/Provider Expectations and Knowledge), a computer adaptive tool grounded in item-response theory that we created, to address this gap by measuring parental and educator knowledge across development domains with precision and speed. This paper details SPEAK's development, including domain construction, cognitive interviewing, expert review, psychometric calibration, and validity evidence. SPEAK offers a flexible, scalable solution for clinical, educational, research, and policy settings. By identifying knowledge gaps, it enables tailored interventions, supports professional development, and informs policy, ultimately improving parent-child interactions and child outcomes. Our tool bridges critical gaps in assessing child development knowledge, advancing research and cross sector collaboration to promote early childhood development worldwide.
With higher education costs consistently outpacing inflation and public funding declining, college affordability has become a critical barrier to economic mobility for middle- and low-income families. While College Savings Accounts (CSAs), or 529 plans, offer tax advantaged vehicles for college savings, their adoption patterns and educational impacts remain poorly understood. Using comprehensive administrative data from over 900,000 Illinois 529 accounts (2000-2023) linked to educational outcomes, plus complementary surveys of account owners and parents, we provide the first large-scale analysis of CSA participation and effectiveness. We find that while CSA adoption has expanded to every ZIP code in Illinois, participation remains concentrated among higher-income, more educated families. Financial literacy emerges as a key barrier: 61% of parents who could save enough to cover half of future college costs still perceive their potential savings as meaningless. Among participants, higher savings are strongly correlated with better educational outcomes, including four-year college enrollment, attendance at selective institutions, and the pursuit of post-graduate degrees. These findings suggest that targeted interventions addressing financial literacy gaps and misperceptions about modest savings could significantly expand CSA effectiveness as a tool for educational equity. Beyond state-level 529 program optimization, our findings suggest several promising avenues for federal policy coordination and institutional innovation.