The phrase “moral economy” comes from two authors. E. P. Thompson (1924-1993) used it in 1971 for the customary expectations that governed what an eighteenth-century English crowd would tolerate from a grain market. A merchant could hold a legal right to charge what the market would bear and still violate the moral economy, because the offense lay in breaching an unwritten understanding about what people owed one another. Lorraine Daston (b. 1951) carried the term into the history of science in 1995, defining a moral economy as a web of affect-saturated values that regulates practice from the inside. Such economies are historically specific, arising and dissolving with the disciplines they govern. And they live in the practitioner, who experiences them as the sense that a finding smells wrong.
Robert Merton (1910-2003) described the scientific ethos in terms of universalism, communalism, disinterestedness, and organized skepticism. Those are rules for evaluating claims once claims are on the table. A moral economy operates earlier. It is a schedule of prices attached to hypotheses. Two explanations can carry comparable empirical support and impose sharply different costs on the scholar who advances them.
Merton’s norms were never the whole sociology of science. Ian Mitroff’s study of the Apollo moon scientists found working counter-norms running alongside them: particularism beside universalism, interestedness beside disinterestedness, emotional commitment beside neutrality, organized dogmatism beside organized skepticism. Scientists carry both codes and do not choose permanently between them. The question is which propositions activate the counter-code.
Contemporary social science has such an economy, and Amy Wax’s 2010 paper on family structure violates it. What produced the hostility was the sense that the paper was doing something illegitimate with the materials.
The central rule it broke was that responsibility flows uphill while causation flows downhill.
When certain people had bad outcomes, explanations that locate the cause in institutions, markets, discrimination, history, neighborhoods, schools, employers, or policy carry a presumption of legitimacy. Explanations that locate any part of the causal chain in genetics, behavior, norms, preferences, family organization, self-command, or cognition face a surcharge. The first class directs attention toward actors with power. The second appears to blame the victim.
The surcharge shows up as a difference in the price of evidence, and Wax’s own review of her field displays it. The marriageable-men hypothesis of William Julius Wilson (b. 1935) attributes the collapse of Black marriage to the disappearance of stable working-class employment. By the estimates of demographers across the methodological spectrum, male earning power and sex-ratio imbalance account for somewhere between five and twenty-five percent of the racial gap, with several estimates far lower. The hypothesis explains a fifth of the phenomenon at its best, and it remains the default account on syllabi. Arline Geronimus’s argument that early nonmarital childbearing constitutes an adaptive response to compressed life expectancy rests on a restricted counterfactual, comparing early nonmarital births to later nonmarital births while never comparing nonmarital to marital, on a sample unrepresentative of the population it describes. It circulated for two decades. Neither survives a week at that standard of fit if the conclusion runs the other way.
The price of evidence has been measured under controlled conditions. Henning Finseraas, Arnfinn Midtbøen, and Kjersti Thorbjørnsrud gave Norwegian social scientists a description of an identical research design on contact between refugees and native-born citizens, randomizing only the reported result. When contact supposedly made attitudes toward immigration more liberal, respondents judged the design better and the research more important than when the same design supposedly made attitudes less liberal. No comparable effect appeared on a nonpoliticized control study about robots. Nothing about the evidence had changed except its implication. The price had.
Scientific racism, eugenics, and theories of inherited inferiority gave behavioral and genetic explanation a bad reputation after the Nazis. Franz Boas (1858-1942) and his students displaced that anthropology by claiming culture as the explanatory unit. Boas produced evidence, including an immigrant anthropometry study covering tens of thousands of measurements whose cranial-plasticity conclusions have been contested. What the moral and political commitments of that circle shaped was the speed of the displacement and the questions the discipline pursued afterward.
The decisive American episode arrived in 1965 with Daniel Patrick Moynihan (1927-2003) and his report on the Black family. William Ryan (1923-2002) supplied the phrase “blaming the victim” that outgrew the controversy that produced it. To describe a behavior that contributed to disadvantage was shifting responsibility from the society that produced the conditions to the people living under them. Wilson wrote later that the controversy made liberal social scientists reluctant to describe behavior that might be regarded as stigmatizing, and that fear of the accusation discouraged investigation of the genetic, cultural and behavioral dimensions of urban poverty.
In 2010, the same year Wax circulated her paper, Mario Luis Small, David Harding, and Michèle Lamont announced in the Annals that culture was back on the poverty agenda. The new cultural sociology analyzed frames, repertoires, identities, perceptions, and meaning-making, in interaction with structural circumstances.
Culture was readmitted under conditions. It could influence how people perceived their options. It could supply repertoires that made some actions easier to imagine. It could mediate responses to structural disadvantage. What stayed hard to say was that some norms and habits produce worse outcomes than others at securing ends that nearly everyone in the population wants.
Wax published into that environment and crossed the line in four places at once. She puts characteristics of actors inside the causal model. She entertains average differences among sociodemographic populations. She describes recurring patterns of conduct as self-defeating even when the people engaged in them are poor. And she argues that traditional moral norms performed useful work because they restricted individual choice.
The fourth is the one that stings. Wax describes pre-1960s sexual mores as heuristics. Clear rules about marriage, fidelity, parenthood, and divorce reduced the foresight and self-command required to produce a stable family. Their removal transferred the burden of regulation from institutions to individuals. People who could supply their own discipline adapted. People who leaned on the external script did not.
That makes the paper more disruptive than a conventional appeal to personal responsibility. Wax asks what arrangements let ordinary people avoid solving the optimization problem at all. Her account of norms runs parallel to behavioral economics. A default enrollment rule gets people to save without requiring them to compute an optimal savings path every payday. A marriage norm does similar work. It sustains a long-term cooperative strategy.
The two paternalisms sit in different places in the moral economy. Automatic retirement enrollment gets called choice architecture and sounds technocratic. Marital fidelity gets called a social norm and sounds moralistic. Steering people away from sugar is public health. Attaching shame to abandoning one’s children is judgment. The behavioral logic is the same in each pair. Only the vocabulary and the price differ. Contemporary social science is full of paternalism. What it will not tolerate is paternalism that speaks in moral language.
A second asymmetry runs through the treatment of individual traits. Dispositional explanations of high-status outcomes are welcome, marketable, and honored. A large body of work on grit, growth mindset, and delay of gratification explains professional and academic success by traits located inside the successful, and it sells to the parents of the successful and gets taught in their schools. Angela Duckworth (b. 1970) and Carol Dweck (b. 1946) built careers on it. The marshmallow paradigm of Walter Mischel (1930-2018) circulated for four decades as a story about what self-command does for a child’s future.
Hold the variable and the measurement constant and point them at the bottom of the distribution, and the reception inverts. Self-command explaining why the children of professors do well is a finding. Self-command explaining part of a group gap is a provocation requiring extraordinary evidence. The methodological critiques of the marshmallow work that arrived after 2018 raised real questions about socioeconomic confounding, and I do not dispute their substance. What tracked the variable’s migration from one use to the other was the energy behind them.
A standard that depends on the status of the population explained rather than on the evidence is a standard of manners. And the manners run in a direction: agency is a compliment paid to people already doing well, and its withdrawal is the courtesy extended to people who are not.
Lee Ross (1942-2021) named the fundamental attribution error in 1977. Observers over-attribute conduct to disposition and under-attribute it to situation, reliably, and worse for people they do not know.
This protects subjects. Findings about populations get used, and the use is rarely gentle. Publish that a group’s members exhibit shorter time horizons and you have produced a sentence that will appear, stripped of qualifications, in the mouth of someone who wants to deny that group a loan.
And it buys access. Ethnography runs on entrée. Kathryn Edin and Maria Kefalas interviewed a hundred and sixty-two single mothers in Philadelphia and produced a book whose empirical value survives its interpretive frame. They could not have done that work while publishing that the central obstacle was the conduct of the men these women chose. The interviews record the women saying so, at length, about infidelity, violence, drug use, and desertion. The book attributes the conduct to bad schools and absent opportunity. The gap between what the subjects said and what the authors concluded is the price of admission, paid in explanatory currency, and every fieldworker knows the exchange rate without having seen it posted.
Mancur Olson (1932-1998) argued that enforcement of a norm is a public good. The benefit diffuses across a population that includes almost nobody the enforcer knows. The cost concentrates on him, in the room, at the time: the awkwardness of judging a colleague’s divorce, the risk of saying in a seminar that children do better with a father at home, the demotion that attaches to anyone who moralizes about sex. Every individual therefore has reason to withhold enforcement while hoping others supply it, and provision settles below what even the enforcers privately believe correct. The same arithmetic governs the scholarly version. Each researcher who declines to publish the uncomfortable finding gains a private benefit and imposes a small diffuse cost on the field’s stock of knowledge. Nobody defects from the economy because defecting is expensive and the returns go to strangers.
Joanna Kempner, Jon Merz, and Charles Bosk (1948-2020) documented how the enforcement runs. They interviewed forty-one academic researchers about knowledge considered too sensitive, dangerous, or taboo to produce, and found that the constraints were largely informal and self-imposed rather than codified. What marked the boundaries were the narrative legacies of past controversies, which working scientists used to locate the lines. Researchers learned where the edges were when their own work or a colleague’s drew rebuke.
Those interviews now have a quantitative successor. Cory Clark and twelve coauthors surveyed 470 American psychology professors about conclusions regarded as taboo within the discipline. The scholars who judged those conclusions more likely to be true reported more self-censorship, not less. Almost all anticipated some social cost from stating their empirical beliefs openly, and tenure did little to remove the fear. Most respondents also opposed suppressing scholarship or punishing scholars merely because conclusions might cause harm.
The combination is the finding. A moral economy does not require a room full of people eager to censor one another. It can be sustained by pluralistic ignorance, with researchers who personally favor open inquiry anticipating sanctions from colleagues who may privately be making the same calculation. Timur Kuran (b. 1954) supplies the name. Preference falsification preserves a public equilibrium that few participants privately endorse, and it holds precisely because each participant reads the silence of the others as evidence of a consensus he alone doubts.
This is how a moral economy should operate. Graduate students observe which questions make advisers uneasy. Scholars learn which formulations require three paragraphs of disclaimer. Referees demand alternative structural accounts with more insistence when the conclusion is uncomfortable. Investigators anticipate the reaction and redesign the project before anyone has to object.
The evidence that the process begins well before peer review arrived in 2026. George Borjas and Nate Breznau examined a many-analyst experiment in which seventy-one teams containing 158 researchers received the same data and the same question about immigration and support for welfare programs, and produced 1,253 regression models. The researchers’ prior views on immigration predicted the direction of their estimates. The route ran through specification. Five ordinary research-design decisions accounted for roughly two-thirds of the difference between the pro- and anti-immigration teams. No fabrication, no falsified analysis, no gatekeeper. Research design had become endogenous to prior belief.
Two features of that result discipline the argument here. Both ideological directions showed the effect, and moderate teams drew better blind-review scores than either extreme. The claim is not that one side is uniquely capable of motivated science. It is that a politically homogeneous field converts ordinary motivated reasoning into a systematic tilt, because the specifications that feel natural to most of the room are the ones that get run.
The costs.
A class of true propositions becomes unpublishable, and a field that cannot publish a proposition cannot test or correct it. Suppression does not produce silence on a question. It produces relocation. The question migrates to writers with no data, no training, no review, and no stake in the population’s welfare, who answer it badly and loudly to an audience that has noticed the professionals declining to speak. Wilson said as much about the years after Moynihan.
Theories with poor fit survive on moral standing. A field that keeps the marriageable-men hypothesis at the center of its account while a fifth of the variance supports it has stopped optimizing for fit, and has lost the ability to tell whether it is making progress.
The economy patronizes. Powerful people get treated as agents. Disadvantaged people get treated as respondents to conditions. The employer discriminates, the landlord excludes, the school fails, the market dislocates, the neighborhood constrains, and the disadvantaged individual adapts. Sometimes that is what happened. When the grammar becomes a default theory of causation, agency itself becomes stratified: people at the top hold enough of it to cause social outcomes, and people at the bottom hold enough to experience outcomes without helping to cause them. Orlando Patterson (b. 1940) has pressed this from a position that makes it hard to wave off, arguing that the refusal to discuss culture and conduct among Black Americans amounts to a denial that they are agents in their own lives.
And morally congenial explanations have victims. Suppose a destructive behavior responds strongly to norms and weakly to the economic variable researchers prefer. A generation spent improving the economic variable while declining to investigate the normative one does not protect the affected population. It withholds a useful explanation from them. Compassion attached to the wrong causal model is not compassionate in its consequences.
The strong version of this argument does not survive every test, and the failures are worth stating. An adversarial collaboration led by Diego Reinero examined 194 psychology studies that later faced replication attempts and found that liberal versus conservative slant predicted neither replication success nor citation frequency nor effect size. What showed modest association with poorer replicability was ideological extremity in either direction. Morally congenial research is not generally false, and the claim here has never needed it to be. The economy changes which questions get asked, how much scrutiny competing explanations receive, and which conclusions impose reputational costs on the person reporting them. Those are effects on the composition of a literature rather than on the truth of any paper in it.
Where the moral economy causes the most damage is in the conflation of explanation with exculpation. Discrimination contributing to an outcome leaves open whether every person experiencing it acted wisely. Behavior contributing to an outcome leaves open whether discrimination is present. A norm proving socially useful leaves open whether those who violate it are wicked. A population differing on average in some trait leaves open where the difference came from, whether it is mutable, and what anyone owes anyone else about it. None of these propositions entails the moral conclusion habitually attached to it.
Wax’s empirical case for population differences in local and global decision-making runs far weaker than her theoretical model. She concedes that the time-discounting literature is messy, that the measurements are inconsistent, and that the evidence linking psychological traits to intimate conduct is thin. That is where to attack the argument. Are the traits measured well? Do the group differences survive controls? Which way does causation run? How much variance do they carry? Can the model outperform its rivals out of sample? Does the normative environment change conduct through the route she specifies?
Those are scientific objections. That hypothesis blames the victim is a statement about the hypothesis’s position in a moral economy.
Researchers choose subjects because some outcomes weigh more with them than others. Societies regulate experiments because knowledge is not worth every price. What a science can demand is that the ledgers stay separate.
The causal ledger asks what produces the outcome. It lets institutions, incentives, history, culture, norms, cognition, personality, choice, inheritance, developmental feedback, and chance compete for explanatory power. No variable earns immunity because its implications are unpleasant, and none earns a presumption because its implications are humane. The moral ledger opens afterward and asks what people owe one another given what turns out to be true.
Separation works because human dignity does not require causal innocence. A poor man need not be a passive product of structures to deserve help. A man who made a disastrous choice keeps his claim on compassion. A group difference is not deserved merely because behavior mediates it. An institution is not innocent merely because individuals act. A society can ask responsibility of people while recognizing that the capacity for responsibility is partly socially produced. Grant those propositions and most of the moral pressure on causal explanation drains away.
Notes
On the term. E. P. Thompson, “The Moral Economy of the English Crowd in the Eighteenth Century,” Past and Present 50 (1971): 76-136. Lorraine Daston, “The Moral Economy of Science,” Osiris 10 (1995): 2-24, at https://www.journals.uchicago.edu/doi/abs/10.1086/368740 and https://www.jstor.org/stable/301910. Daston’s definition of a moral economy as a web of affect-saturated values is the sentence the essay leans on, and her insistence that these economies are historically specific is what makes the present one a subject rather than a scandal. For the concept’s wider career, Norbert Götz, “‘Moral Economy’: Its Conceptual History and Analytical Prospects,” at https://www.tandfonline.com/doi/full/10.1080/17449626.2015.1054556. Robert Merton’s four norms come from “The Normative Structure of Science” (1942), reprinted in The Sociology of Science (1973). For the counter-norms that run alongside them, Ian I. Mitroff, “Norms and Counter-Norms in a Select Group of the Apollo Moon Scientists: A Case Study of the Ambivalence of Scientists,” American Sociological Review 39 (1974): 579-595, at https://www.andreasaltelli.eu/file/repository/mitroff_OCR.pdf.
On the Wax paper and its field review. Amy L. Wax, “Diverging Family Structure and ‘Rational’ Behavior: The Decline in Marriage as a Disorder of Choice,” University of Pennsylvania Public Law and Legal Theory Research Paper No. 10-17, at https://ssrn.com/abstract=1592424. Her demolition of the structural accounts runs through the paper’s third section, and her own summary of the demographic estimates on marriageable men appears at notes 30 and 31. The Wilson hypothesis is stated in The Truly Disadvantaged (1987), whose opening chapter also carries Wilson’s account of the professional silence that followed the Moynihan report. Kathryn Edin and Maria Kefalas, Promises I Can Keep (2005), is the ethnography discussed above.
On the price of evidence, measured. Henning Finseraas, Arnfinn H. Midtbøen, and Kjersti Thorbjørnsrud, “Ideological Biases in Research Evaluations? The Case of Research on Majority-Minority Relations,” Scandinavian Political Studies 45 (2022): 370-381, at https://onlinelibrary.wiley.com/doi/full/10.1111/1467-9477.12229, is the preregistered experiment in which the conclusion was randomized and the design held constant. George J. Borjas and Nate Breznau, “Ideological Bias in the Production of Research Findings,” Science Advances 12 (2026): eadz7173, at https://pmc.ncbi.nlm.nih.gov/articles/PMC12757037/ and https://pubmed.ncbi.nlm.nih.gov/41477854/, supplies the many-analyst result. Note that both papers concern immigration rather than family structure, and that neither observes the specific economy this essay describes. What they establish is that the pricing effect is real and measurable in an adjacent domain.
On how the boundaries get enforced. Joanna Kempner, Jon F. Merz, and Charles L. Bosk, “Forbidden Knowledge: Public Controversy and the Production of Nonknowledge,” Sociological Forum 26 (2011): 475-500, at https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1573-7861.2011.01259.x, with an open copy here. The earlier and shorter statement is Kempner, Clifford Perlis, and Merz, “Forbidden Knowledge,” Science 307 (2005): 854, at https://science.sciencemag.org/content/307/5711/854. Their finding that most constraints are informal or self-imposed, and that memories of past controversies mark the lines, is the empirical backbone of the argument here. The quantitative successor is Cory J. Clark and colleagues, “Taboos and Self-Censorship Among U.S. Psychology Professors,” Perspectives on Psychological Science 20 (2025): 941-957, at https://pmc.ncbi.nlm.nih.gov/articles/PMC12408927/ and https://journals.sagepub.com/doi/10.1177/17456916241252085, whose crucial result is that belief in a taboo proposition predicts more self-censorship rather than less. For the collective-action engine, Mancur Olson, The Logic of Collective Action (1965). For the equilibrium it sustains, Timur Kuran, Private Truths, Public Lies: The Social Consequences of Preference Falsification (Harvard University Press, 1995).
On the trait literature and its asymmetric reception. Walter Mischel, Yuichi Shoda, and Philip Peake’s original delay-of-gratification papers ran in Science (1989) and Developmental Psychology (1990). The conceptual replication is Tyler Watts, Greg Duncan, and Haonan Quan, “Revisiting the Marshmallow Test,” Psychological Science 29 (2018): 1159-1177, at https://journals.sagepub.com/doi/abs/10.1177/0956797618761661, with an open copy here. Their bivariate correlation ran half the size of the original and fell by two thirds under controls for family background, early cognitive ability, and home environment. The exchange that followed is worth reading in full, since it shows the field arguing about controls rather than about politics: Armin Falk, Fabian Kosse, and Pia Pinger, “Re-Revisiting the Marshmallow Test,” here with the Watts and Duncan reply on confounding and construct clarity. Lee Ross’s “The Intuitive Psychologist and His Shortcomings” (1977) is where the attribution error gets its name.
On what the argument does not establish. Diego A. Reinero and colleagues, “Is the Political Slant of Psychology Research Related to Scientific Replicability?,” Perspectives on Psychological Science 15 (2020): 1310-1328, at https://journals.sagepub.com/doi/10.1177/1745691620924463 and https://pubmed.ncbi.nlm.nih.gov/32812848/, is an adversarial collaboration across 194 studies subjected to replication attempts, and it finds no relation between ideological direction and replication success, citation, or effect size. Anyone tempted to read this essay as a claim that congenial findings are false should read Reinero first. The Borjas and Breznau result runs the same direction, since both ideological camps showed the specification effect and moderate teams drew better blind-review scores.
On Boas. The immigrant anthropometry study is Changes in Bodily Form of Descendants of Immigrants (1912). The reanalysis dispute runs through Corey Sparks and Richard Jantz in the Proceedings of the National Academy of Sciences (2002) and Clarence Gravlee, H. Russell Bernard, and William Leonard in American Anthropologist (2003), the latter concluding that Boas’s central plasticity finding largely held. The claim in the text is about the speed of a disciplinary displacement and the questions pursued afterward, not about an absence of evidence.
On the Moynihan aftermath. Daniel Patrick Moynihan, The Negro Family: The Case for National Action (Office of Policy Planning and Research, U.S. Department of Labor, 1965). William Ryan, Blaming the Victim (1971), supplied the phrase. Readers who want the controversy reconstructed rather than summarized should start with James Patterson, Freedom Is Not Enough (2010).
On culture readmitted under conditions. Mario Luis Small, David J. Harding, and Michèle Lamont, “Reconsidering Culture and Poverty,” Annals of the American Academy of Political and Social Science 629 (May 2010): 6-27, at https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2131376, with the Harvard open copy at https://scholar.harvard.edu/files/lamont/files/reconsidering_culture_and_poverty.pdf and the full volume at https://archive.org/details/reconsideringcul0629unse. The volume includes a contribution from Wilson. Note the timing: this appeared the same season Wax circulated her draft, which makes the pair a natural comparison for anyone who wants to see the permitted and the forbidden versions of the same move side by side.
Further reading. Orlando Patterson, The Ordeal of Integration (1997) and Rituals of Blood (1998), argue the agency point from inside the sociology of race; his 2006 New York Times essay “A Poverty of the Mind” states it compactly. On the demographic side of the family argument, Andrew Cherlin, The Marriage-Go-Round (2009), Charles Murray, Coming Apart (2012), Robert Putnam, Our Kids (2015), and Melissa Kearney, The Two-Parent Privilege (2023). Philip Tetlock’s work on sacred values and taboo tradeoffs supplies a psychology for the reaction described here, and readers who want it should start with “The Psychology of the Unthinkable” (2000). For a companion piece on who selects a normative regime and who bears its costs, see the essay on status hierarchy and normative selection that shares this argument’s premises.
Disruptive Papers
The strongest of these dissolve a morally convenient dichotomy rather than delivering an unwelcome result. A reader cannot neutralize them by switching political teams, which is what makes them expensive to price.
Augustine Kong and colleagues, “The Nature of Nurture: Effects of Parental Genotypes,” Science 359 (2018): 424-428. The design exploits alleles parents carry but do not transmit. Since the child never inherited them, any association with the child’s educational attainment cannot be a direct genetic effect on the child. A polygenic score built from nontransmitted parental alleles predicted the child’s educational attainment at 29.9 percent the strength of the transmitted score. They named it genetic nurture. The partition between heredity and environment stops being clean in both directions at once. A variable coded environmental can carry genetic covariance, and a genetic estimate can carry parental behavior.
Laurence Howe and colleagues, “Within-Sibship Genome-Wide Association Analyses Decrease Bias in Estimates of Direct Genetic Effects,” Nature Genetics 54 (2022): 581-592, free at PMC. 178,086 siblings across 19 cohorts. Comparing siblings strips out population stratification, assortative mating, and family-level genetic nurture. Estimated associations shrank substantially, roughly forty-seven percent for educational attainment, and many survived. Neither camp gets what it wants. Direct inheritance, indirect genetic effects, and social environment remain entangled, and nobody is granted causal innocence.
Aaron Chalfin, Benjamin Hansen, Emily Weisburst, and Morgan Williams Jr., “Police Force Size and Civilian Race,” American Economic Review: Insights 4 (2022): 139-158, free at PMC. 242 cities, 1981 to 2018, two instrumental-variable strategies. An additional officer prevents about a tenth of a homicide, and the per-capita homicide benefit runs roughly twice as large for Black residents as for white residents. More police also means fewer serious-crime arrests and more low-level arrests, with the second burden falling disproportionately on Black Americans. Both slogans about policing turn out true along different margins. The finding refuses to announce the researcher’s politics, which is the sorting device the moral economy runs on.
Michael Schaerer and colleagues, “On the Trajectory of Discrimination: A Meta-Analysis and Forecasting Survey Capturing 44 Years of Field Experiments on Gender and Hiring Decisions,” Organizational Behavior and Human Decision Processes 179 (2023): 104280. Preregistered, 85 field audits, 244 effects, 361,645 applications, an independent red team, and a forecasting component. Discrimination against women in male-typed and mixed jobs declined and reversed slightly after 2009. Discrimination against men in female-typed jobs persisted. Then the second catch: scientists correctly anticipated improvement for women, overestimated how much anti-female discrimination remained, and did not anticipate the anti-male effect at all. The paper measures the phenomenon and the field’s expectations about it in the same design.
The next group holds a design constant and gets an answer the field did not want.
Valentin Bolotnyy and Natalia Emanuel, “Why Do Women Earn Less than Men? Evidence from Bus and Train Operators,” Journal of Labor Economics 40 (2022): 283-323, PDF at https://scholar.harvard.edu/files/bolotnyy/files/be_gendergap.pdf. One unionized employer, identical posted wages, seniority-driven promotion, and still eighty-nine cents on the dollar, arising from 1.5 fewer overtime hours and 1.3 more unpaid hours off per week. Occupational sorting, managerial discretion, and negotiation are removed by construction.
Cody Cook, Rebecca Diamond, Jonathan Hall, John List, and Paul Oyer, “The Gender Earnings Gap in the Gig Economy: Evidence from over a Million Rideshare Drivers,” Review of Economic Studies 88 (2021): 2210-2238, PDF at https://web.stanford.edu/~diamondr/UberPayGap.pdf, working version at NBER 24732. More than a million workers, a sex-blind algorithm setting pay and dispatch, and a seven percent hourly gap accounted for by experience, where and when people drove, and speed. Customer discrimination explained none of it. The result does not show that gaps elsewhere are preference-driven. It shows that a substantial gap can appear where the standard suspects are absent.
Peter Arcidiacono, Josh Kinsler, and Tyler Ransom on Harvard admissions, principally “Asian American Discrimination in Harvard Admissions,” European Economic Review 144 (2022): 104079, and “Legacy and Athlete Preferences at Harvard,” Journal of Labor Economics 40 (2022): 133-155, with everything collected at https://psarcidi.github.io/. Two competent teams working the same applicant-level records, with David Card running the opposing analysis for Harvard. That is the rarest thing in social science and it lets a reader watch specification choice operate in daylight. The second paper is the one that keeps it from being a partisan document: more than forty-three percent of white admits were legacy, athlete, dean’s list, or faculty children, and roughly three quarters of those would have been rejected otherwise.
Lincoln Quillian, Devah Pager (1972-2018), Ole Hexel, and Arnfinn Midtbøen, “Meta-analysis of field experiments shows no change in racial discrimination in hiring over time,” PNAS 114 (2017): 10870-10875, free at PMC. Twenty-eight audits, 55,842 applications, whites receiving thirty-six percent more callbacks, no decline since 1989. This one points the other way and belongs on any list that expects to be taken seriously. Correspondence audits also have the cleanest identification in the discrimination literature, which is precisely why the result is hard to price down.
The last group is about the machinery of belief rather than about any social outcome.
Patrick Forscher, Calvin Lai, Jordan Axt, Charles Ebersole, Michelle Herman, Patricia Devine, and Brian Nosek, “A Meta-Analysis of Procedures to Change Implicit Measures,” Journal of Personality and Social Psychology 117 (2019): 522-559, preprint at https://osf.io/preprints/psyarxiv/dv8tu/. 492 studies, 87,418 participants. Implicit measures move, effects run small, behavioral effects run trivial, and changes in implicit measures do not mediate changes in behavior. Devine and Nosek built the field. The paper does not deny implicit attitudes. It removes the inference that moving a score is a demonstrated route to changing conduct, which is the inference the training industry sells.
Jeffrey Dahlke, Paul Sackett, and Nathan Kuncel, “Effects of Range Restriction and Criterion Contamination on Differential Validity of the SAT by Race/Ethnicity and Sex,” Journal of Applied Psychology 104 (2019): 814-831. 363,004 students at 107 institutions. Correcting for admissions-induced range restriction and for students taking different courses, first-year predictive validity for Black and Hispanic students was not significantly different from validity for white students, with four-year outcomes and other comparisons more complicated. The disruptive proposition is narrow and load-bearing: a group difference in average scores is not itself evidence that the instrument is biased. Disparity and measurement bias are separate empirical claims.
George Borjas and Nate Breznau, “Ideological Bias in the Production of Research Findings,” Science Advances 12 (2026), eadz7173, record at PubMed. Seventy-one teams, 158 researchers, identical data and question, 1,253 models. Prior views predicted estimates, and five ordinary design choices carried about two thirds of the difference between camps. Both directions showed it and moderate teams drew better blind-review scores. No fraud, no censorship, no gatekeeper. The dependent variable here is science, which makes it the closest thing to an experimental demonstration of a price of evidence.
Cory Clark and colleagues, “Taboos and Self-Censorship Among U.S. Psychology Professors,” Perspectives on Psychological Science 20 (2025): 941-957, belongs alongside these as a companion rather than a peer. It is a survey and carries no causal design. Its value is the finding that professors who judged taboo propositions more likely to be true reported more self-censorship, which is the shape a pluralistic-ignorance equilibrium leaves in survey data.
What happens on the inside when an elite academic who regards himself as a disinterested servant of the truth reads a disruptive paper?
Claude says:
Nobody has written the scene you want. Three things come close.
Robert Putnam (b. 1941) wrote the nonfiction version about himself. His 2007 paper on ethnic diversity and social capital found that diversity depressed trust in the short run, including trust within groups, and he said in print that he had delayed publication while looking for countervailing evidence and had wanted to be able to report a remedy alongside the finding. That is a man describing his own hesitation, in a journal, with the delay itself as part of the record.
Alice Dreger’s Galileo’s Middle Finger, 2015, is the memoir of watching this happen from inside and then deciding to become an inconvenience about it.
For fiction, Allegra Goodman’s Intuition, 2006, is the closest thing that exists. A postdoc’s result is too good, and the novel is about who wants it to be true, who is permitted to doubt it, and what doubting costs. C. P. Snow (1905-1980) is the ancestor, particularly The Affair, where a college decides what it can afford to have happened.
SCENE:
The paper comes in from a former student on a Tuesday, with no message, only the subject line: thought you’d want to see this. He opens it between the eleven o’clock and the department lunch. He reads the abstract twice.
The first thing he does is scroll to the affiliations. This is not a political act. He has done it for thirty years, because a name and a department tell him how much weight the standard errors will bear, and because he can save himself forty minutes by learning that the third author trained under someone whose work he already distrusts. The affiliations are good. Better than good. One of them is a place he has sent his own students.
So he reads the data section.
He reads it the way he reads a referee assignment, which is to say he reads it looking for the place where it breaks. This too is not political. It is the only way anyone reads anything in his field. He finds three candidates in the first pass. The sample is drawn from administrative records that exclude a population he happens to know is excluded in a nonrandom way. The controls include a variable that is plausibly post-treatment. And the main specification uses a fixed effect that will absorb more variation than the authors seem to want it to.
He writes these down on the back of a seminar announcement. He feels, writing them, the small clean pleasure of competence. He is good at this. He was good at it when he was twenty-eight and he is better now.
Then he goes to the appendix, because the appendix is where a serious team puts the robustness checks, and he wants to know whether they anticipated him.
They anticipated him. All three.
The exclusion is addressed in Table A4, with bounds. The post-treatment control is dropped in column five and the coefficient moves by four percent. The fixed effect is relaxed in three ways across two pages, and the result survives all of them, with the confidence interval widening in the direction he would have predicted and not far enough to matter.
He sits back.
What happens next takes about four seconds and he will not remember it by dinner.
He thinks: the effect size still seems large. He does not think this because he has a prior about the effect size. He has no prior about the effect size. He has never worked on this outcome. He thinks it because the number, if true, would require him to say something in a seminar that he does not want to say in a seminar, and the not-wanting has arrived before the reasoning, dressed as the reasoning, and he has no instrument that can tell the two apart.
So he looks for a fourth problem.
He finds one. The outcome is self-reported in the second wave. This is a real limitation and he is right to note it. He notes it. He notes it with more energy than he noted the first three, and he does not observe the difference in energy, because energy is not something he has ever been trained to measure in himself.
He looks at the clock. He is late for lunch.
Walking down the corridor he composes, without deciding to, the sentence he will use if the paper comes up. He is not going to attack it. He would never attack it. He is going to say that it is careful work and that he would want to see it replicated on a different population before drawing conclusions. Every word of that is true. He would want that. He wants that about everything.
He has never once said it about a paper whose result ran the other way.
At lunch a colleague mentions the paper. There is a pause of the kind that is not long enough to be called a pause. Someone says the authors are serious people. Someone else says he heard the identification is fragile, and the man who read the appendix that morning, who knows the identification is not fragile, who has the seminar announcement in his jacket pocket with three refuted objections written on the back of it, says nothing at all.
He says nothing because the room has not asked him anything. He tells himself this on the walk back, and it is true, and he is a scrupulous man, and it is the last time he thinks about the paper for six weeks.