Andrew Gelman (b. 1965) answered two of my emails this year.
On April 6, 2026, I emailed him about the unwritten rules governing speech on elite campuses, and asked whether a passage I had written about Columbia struck him as fair:
Columbia is the most volatile campus, and its no-go zones shift faster than anywhere else because two powerful coalitions are in open conflict. High-intensity activist networks and an administration under significant federal pressure collide constantly, and the boundaries move with each news cycle. Pro-Israel speech in activist spaces requires heavy hedging. Pro-Palestinian speech that crosses new procedural lines installed under federal scrutiny carries its own risks. The deeper prohibition is being legible to neither coalition: floating above the conflict reads as moral evasion, and the system punishes that more reliably than it punishes taking either side. Columbia students face both peer friction and administrative friction, which produces the lowest free speech scores in the country. The tacit rule is that you must pick a side or perform neutrality with great care. (source)
He replied: “Hi, I have no idea what is meant by a no-go zone. I have not seen any prohibitions myself, but I have read about some things such as Barnard not letting students put signs on their dorm room walls.”
I meant the unwritten rules about what you can and cannot say publicly without social penalty, self-censorship driven by fear of consequence.
He answered: “In that case, I know of no such unwritten rules.”
Four months later, on August 26, 2026, he posted two other questions I had sent him, about whether the culture of research has changed since the replication crisis and about how institutions lose the ability to perceive their own situation. He answered both at length. Then he added this:
Ford posted this discussion on his blog. It was kinda weird seeing myself discussed as a sociological object, but, fair enough, I’m a public figure, and people can say what they want as long as they don’t misrepresent my writings or claim that I said something I never said.
Ford’s assessment is accurate that I’m not very good at strategic behavior so often I don’t even try. It’s similar to how I’m a bad negotiator so usually I’ll just try to make my goals clear and not try to optimize, following the “Getting to Yes” principle that the main thing getting in the way of smooth negotiation is ignorance of other people’s goals. I think back to various successful and botched negotiations I’ve been involved with in the past, and almost always the problems come with struggles over details without there being clarity on the goals of the different parties.
Put the two exchanges together. In the second, a man reads a sociological account of himself, grants its central claim about his own character, and files no objection. In the first, the same man reports that a sociological category offered to him corresponds to nothing he can see. The distance between the responses is the subject of this essay.
The record
Gelman was born in Philadelphia into a family with intellectual range. His sister Susan Gelman became a developmental psychologist. His uncle Woody Gelman was a cartoonist. He went to MIT as a National Merit Scholar, took degrees in mathematics and physics in the mid-1980s, then moved to Harvard for graduate work in statistics, finishing his doctorate in 1990 under Donald Rubin, whose work on missing data and causal inference had already reshaped how empirical researchers thought about what they could learn from observational studies. Gelman’s dissertation, Topics in Image Reconstruction for Emission Tomography, took up a problem in medical imaging. A positron emission tomography scanner records noisy, indirect photon counts, and the statistician has to infer from them the distribution of activity inside a living brain. Rubin suggested the topic. Gelman did not invent the field he entered. Shepp and Vardi had published the maximum-likelihood approach to emission tomography in 1982, Stuart and Donald Geman had published on Bayesian image restoration in 1984, and his third chapter presents itself as working inside that literature. What he marks as his own are the bias and variance expressions for regional averages in Chapter 7, worked into computable form rather than left as matrix algebra. His own verdict, given later, is that the thesis was never published whole, that he mailed out perhaps a hundred copies, and that most of it served as an education for him.
Read now, it is the man in outline. The organizing problem is one where the data cannot settle the question. Limited resolution and incomplete sampling mean there is no assumption-free estimator waiting to be found, and he says so twice, in the abstract and in the conclusion: every image estimate must rest on assumptions that cannot be verified from the data. His response is to name the assumptions and study what happens when they fail. He adds that estimation bias tends to show up as a blur. The blur is the honest width of what the instrument can see, and the wide interval he later published when the market paid for clean results is the same object.
Chapter 5 has the ancestor of partial pooling. Conditional autoregressive priors let neighboring pixels inform one another, which is the maneuver the later hierarchical work performs with states and schools and demographic cells. Chapter 8 has the ancestor of design analysis. He lays out region effects, region by task effects, region by task by patient effects, and their replications as nested variance components, then uses the variance expressions to ask how many patients and tasks a future study would need to detect an effect of a given size. And near the end he refuses to name a best reconstruction, on the ground that the virtues of a method depend on the intended use of the image. A reconstruction good for looking at a brain may be poor for estimating averages within anatomical regions. Thirty-six years of arguing against judging a procedure by an abstract criterion detached from its application start there.
The PET study behind all this had five subjects performing eye-movement tasks. Gelman writes in the thesis that the sample was too small to estimate the effects of interest. He does not discard it. He turns it into evidence about measurement error, reconstruction bias, and variation between patients, and uses it to work out what a better study would require. He also reports, in his own conclusion, that a direct implementation of the maximum likelihood estimate, which he does not present, indicates that the standard statistical model of PET data is incomplete in practice. He ran it, it failed, and he told the reader in a document nobody was going to check. Two years later he ran the test properly in Statistical Analysis of a Medical Imaging Experiment, where the chi-squared check confirms the lack of fit. The companion paper from the same year, Who Needs Data? Restricting Image Models by Pure Thought, rules out models by working through implications they cannot survive, including what happens to their conclusions under arbitrary rescaling. That is the piranha problem in its infancy.
What is absent matters as much. There is no garden of forking paths, no Type S and Type M errors, no publication bias, nothing about significance testing, no causal inference in the later sense, and no political science. The computation is EM and posterior modes rather than the Markov chain machinery of the textbooks. The polemical voice is nowhere; the prose is sober and conventional. The critic arrives later. What runs continuously from 1990 is the epistemology.
On May 16, 2010, Gelman wrote:
Around the time I was finishing up my Ph.D. thesis, I was trying to come up with a good title–something more grabby than “Topics in Image Reconstruction for Emission Tomography”–and one of the other students said: How about something like, Female Mass Murderers: Babes Behind Bars? That sounded good to me, and I was all set to use it. I had a plan: I’d first submit the one the boring title–that’s how it would be recorded in all the official paperwork–but then at the last minute I’d substitute in the new title page before submitting to the library. (This was in the days of hard copies.) Nobody would look at the time, then later on, if anyone went into the library to find my thesis, they’d have a pleasant surprise. Anyway, as I said, I was all set to do this, but a friend warned me off. He said that at some point, someone might find it, and the rumor would spread that I’m a sexist pig. So I didn’t.
He joined Berkeley’s statistics faculty in 1990, moved to Columbia in 1996, and stayed, accumulating appointments in both statistics and political science, directing the Applied Statistics Center, and in 2017 taking the Higgins Professorship. A technical statistician turned into a public epistemologist, an institutional auditor, a methodological conscience for the social sciences, and a blogger whose daily output shaped the research culture of a generation.
The technical work came first. Bayesian Data Analysis, written with John Carlin, Hal Stern, David Dunson, Aki Vehtari, and originally Rubin, is in its third edition and serves as the standard reference for applied Bayesian reasoning across statistics, epidemiology, social science, and machine learning. Its distinction is the insistence on workflow. Bayesian analysis is an iterative process of model building, checking, revision, and expansion. Posterior predictive checks, the technique of asking whether data simulated from your fitted model resemble the data you saw, become in Gelman’s treatment the core of honest inference. If your model cannot generate data that looks like what you observed, the model is wrong somewhere and you should find out where.
That emphasis on checking connects to his work on multilevel modeling, laid out in Data Analysis Using Regression and Multilevel/Hierarchical Models with Jennifer Hill and its successor Regression and Other Stories with Hill and Vehtari. Multilevel models let parameters vary across groups while borrowing strength across them through partial pooling. They existed before Gelman. He made them accessible and put them at the center. Complete pooling treats all groups as identical. No pooling treats each as an island. Partial pooling holds that groups share something while differing in ways the data can inform. The move is technically superior and epistemically humble at once. It encodes the belief that your data know something about the world and not everything.
The political science work grew alongside. Red State, Blue State, Rich State, Poor State, with David Park, Boris Shor, and Jeronimo Cortina, took a puzzle that political commentary kept getting wrong. Red states were poor and blue states rich at the state level, and the pattern inverted at the individual level: within any given state, higher-income voters leaned Republican. The paradox dissolves once you model individual voting within states. The book shows what happens when you analyze data at the right level of aggregation.
The forecasting work extended the approach. His collaboration with The Economist on presidential models in 2020 and 2024 stands out for what it declined to do. The models produced wide intervals. They combined economic fundamentals with polling adjusted for non-sampling error. They treated a forecast as a posterior distribution over outcomes. In a market where forecasters competed on the display of confidence, his models described the uncertainty as it stood, which made them a worse product and a better description.
The work on redistricting, conducted with Gary King, helped give quantitative form to arguments about electoral fairness. The standard of partisan symmetry, developed by King with Bernard Grofman and applied in this line of research, holds that a map is fair if both parties receive the same seat share at the same vote share. That translated a moral intuition into a measurable property of electoral systems. The redistricting research also produced a finding advocates on both sides preferred to skip: gerrymandering often fails to deliver the partisan effects it intends, because the political environment carries enough uncertainty that mapmakers cannot engineer the outcomes they want.
On the rationality of voting, Gelman worked with Aaron Edlin and Noah Kaplan on the social-preferences argument, and later with Nate Silver and Edlin on the probability that a vote proves decisive. The puzzle in rational choice theory is familiar. If the chance your vote decides an election is vanishingly small, why vote? The standard answer treats voting as expressive. The Edlin, Gelman, and Kaplan answer works the expected utility calculation through. If voters hold social preferences, and the benefit of a preferred candidate winning multiplies across the whole affected population, the expected utility of voting stays substantial even in a large electorate. The argument against voting’s rationality rests on a contestable assumption about the scope of what voters want.
Any of this would have made a career. What sets Gelman apart is that it ran alongside a deepening preoccupation with how knowledge gets produced, with what researchers are doing when they claim to do science, and with the gap between the methods described in papers and the methods used in labs.
The garden of forking paths, named for the Borges story, is the clearest expression of that preoccupation. Jorge Luis Borges (1899-1986) supplied the image. The phenomenon is the ordinary condition of modern empirical research. A researcher collects data, looks at it, and makes a series of reasonable decisions: which outliers to drop, which covariates to include, which subgroups to examine, which outcome to feature. Each choice is defensible. Together they guarantee that something significant emerges. The space of possible analyses is large, the researcher walks it in real time in response to what the data show, and the resulting p-value does not mean what p-values are supposed to mean. A p-value tells you the probability of data as extreme as yours under the null, given that you committed to the analysis before you looked. It tells you nothing once the data shaped the analysis.
The point is sociological as much as statistical. Researchers are trained to explore their data, to hunt patterns, to understand their measurements before settling on a final model. That exploration serves understanding and kills the inference the published p-value is supposed to carry. Gelman’s contribution was to name the problem and refuse to treat it as a rare pathology.
Type S and Type M errors follow. When an underpowered study reaches significance, the result exaggerates. The true effect, if any, runs smaller than what was measured, and there is a real chance the sign is wrong. Type S errors point in the wrong direction. Type M errors inflate magnitude. Neither is random. Both fall out of the publication process in low-powered research environments. Gelman and his collaborators worked the mathematics through, showing that in fields where effects run small and studies run underpowered, the published literature will be dominated by exaggerated and unreliable estimates with no fraud anywhere in the chain.
The replication crisis that surfaced between 2011 and 2015 bore the analysis out. Celebrated studies across psychology, social science, nutrition, and medicine failed to replicate at rates that follow from what was known about power and forking paths, and that shocked people who had not thought about what the publication process produces. The power posing study by Amy Cuddy and Dana Carney, claiming that expansive postures raise testosterone, lower cortisol, and improve risk tolerance, was the most prominent casualty. The sample was forty-two people. The reporting selected among many measured hormonal and behavioral outcomes. The finding generated a TED talk viewed tens of millions of times and a career built on its emotional resonance. Gelman’s critique stayed on the methods: the sample was too small, the selective reporting was a textbook walk through the garden, and the claimed effect sizes were implausibly large against the noise level of the study.
Cuddy’s defenders answered. A piece in the New York Times Magazine framed the episode as scientific bullying and charged Gelman with cruelty and a failure of collegiality. A proper critic, the argument ran, would have reached out privately, helped the researcher fix her errors, and spared her the humiliation. Gelman’s answer stayed consistent. Science is a public activity. A claim published in a peer-reviewed journal and carried to millions through a TED talk is not a private matter between two researchers. The obligation to defend published claims in public is not cruelty. The practice of helping famous researchers repair their errors in private, shielded from scrutiny, is a form of corruption. It protects reputations at the cost of the public record.
The Brian Wansink case made the argument harder to duck. Wansink directed Cornell’s Food and Brand Lab and had built a prolific career on research about food behavior, portion size, and the environmental determinants of eating. A blog post in which he praised a student for squeezing four papers out of a dataset that had failed to produce significance triggered a systematic examination of his published work. Gelman and others, including Jordan Anaya and Nick Brown, found impossible numbers, inconsistencies across papers reporting the same studies, and patterns consistent with pervasive p-hacking. More than fifteen of his papers were retracted. Cornell investigated, found academic misconduct, and Wansink resigned. Post-publication scrutiny worked when somebody pursued it.
These fights made Gelman a polarizing figure. Susan Fiske circulated a draft essay describing critics like him as methodological terrorists. The phrase crystallized a real tension. The old system of quality control ran through pre-publication peer review, conducted in private by a small number of experts who shared the professional norms and the social networks of the authors they reviewed. Conducting critique on a blog anyone can read violates those norms. It puts methodological arguments that used to travel in private letters between specialists in front of journalists, policy advocates, and the public. It strips the protection that professional insularity gave to prominent work that could not survive examination.
Defending that insularity misses what was at stake. The old system had failed. Private collegial correction let the garden of forking paths flourish for decades and produced a published literature in social psychology and nutrition science that could not be trusted.
The blog, Statistical Modeling, Causal Inference, and Social Science, which he has run since 2004, is where most of this played out. It is a daily practice of reasoning in public, a running commentary on the methods and findings of empirical research across many fields, and a space where technical arguments meet an informed and hostile readership. Posts range from statistical critiques of individual papers to reflections on research culture, teaching, politics, literature. The comment section at its best works as extended peer review of Gelman’s own arguments, with methodologically sophisticated readers from many fields pushing back and correcting in real time.
Its influence is large. It shaped how a generation thinks about uncertainty, model checking, and the ethics of publishing. It gave a public vocabulary to problems that had no name. The garden of forking paths, researcher degrees of freedom, Type S and Type M errors, the time-reversal heuristic, all circulate across disciplines partly because they were named in a form non-specialists could carry. The blog also changed what was socially possible. By showing that serious technical criticism could run in public without the cover of anonymous review, it normalized a form of accountability the old system had made nearly impossible. Stephen Turner describes what the blogosphere does at its best: it aggregates testimony the specialist channels filter out, performs folk sociology on expert claims, supplies factual material that qualifies expert assertions, and forces experts to justify themselves. The Wansink correction ran through every one of those steps.
The reform agenda he pushed for two decades is mainstream enough now that people forget how contested it was. Pre-registration, open data, post-publication review, the retirement of p less than 0.05 as a criterion for publishability, the treatment of replication as a basic activity: methodologists accept these, journals encode them, funders require them. Gelman did not do this alone. The public argument he conducted through the blog, the papers, the books, and the confrontations helped produce the conditions in which the changes became possible.
The contribution to Stan, the probabilistic programming language built with Bob Carpenter and others, is a different contribution. Stan puts Bayesian computation in reach of researchers who cannot derive sampling algorithms from scratch. It handles the hard work of exploring posterior distributions and leaves the researcher to specify and interpret the model. This is infrastructure. It lowered the barrier to entry across an enormous range of fields, and the research that has flowed through it is substantial.
The multilevel regression and poststratification work, developed with Yair Ghitza and others, is now standard for small-area estimation in political science and survey research. Combine a large national survey with census data on demographic composition and you can estimate opinion at the state or district level from a handful of local respondents. The model borrows strength across demographic groups and geographies. The poststratification step reweights predictions to the composition of each area. Accurate subnational estimates become achievable at a fraction of what state-level surveys used to cost.
The Xbox study, conducted with data from the Microsoft gaming platform during the 2012 election, pushed the approach as far as it goes. The Xbox sample skewed heavily toward young men. After adjustment, the estimates tracked traditional probability-based polls. The demonstration unsettled people because it suggested that representativeness at the sampling stage matters less than assumed, and that adjustment can pull a reliable signal out of biased data. That has practical consequences for survey design, which had treated probability sampling as a necessary condition for valid inference.
The work on polling variance, again with King, addressed a question that generated years of confusion in political journalism. Presidential polls swing week to week during a campaign, and the final outcome is often predictable months out from economic and political fundamentals. Why does the polling move so much if the result is largely set? Their answer was the enlightened preferences hypothesis. Voters hold underlying preferences shaped by demographics, ideology, and their read on economic conditions. The campaign functions as an information environment that helps voters converge on those preferences. Early polls measure noise around a stable signal. As the election nears and information improves, the noise falls away and the polls settle toward the fundamentals.
Political journalists ignored the finding, since it implied that most of what they cover as consequential, the daily movement driven by speeches, gaffes, debates, and events, is noise around a signal fixed before the campaign began. Campaigns still do something. They do a good deal less than horse-race journalism assumes.
The teaching work, in Teaching Statistics: A Bag of Tricks with Deborah Nolan and Active Statistics with Vehtari, carries the same commitments into pedagogy. The books run on stories, real data, and the link between a method and the question it answers. They resist teaching statistics as a set of procedures to be executed and graded. Statistics is a way of reasoning about uncertainty, and students learn it by doing it.
His March 2026 reply to a question about whether research culture has reformed shows where his assessment now stands, and it separates two questions people run together. Inside academic science there has been reform: more skepticism toward underpowered studies, more pre-registration, more open data, less tolerance for the crude p-hacking of the pre-crisis era. The public intellectual ecosystem that science feeds has not reformed. The pipeline that used to run from researchers through journals to NPR, TED, and Gladwell-style popularization has been bypassed. Andrew Huberman and the direct-to-audience health industry need no academic credentialing to reach millions. They produce and distribute junk directly. The reform of academic science arrived alongside the collapse of the infrastructure that made academic science consequential for public belief.
That is a sober reading from a man who spent decades improving that infrastructure from inside. It suggests the replication crisis was a symptom of something deeper. The relationship between research and public knowledge has been disrupted in ways that better methods inside universities cannot repair.
His remark about Columbia, in the same reply, applies a structural analysis to university governance. Universities have an executive function and thin legislative and judicial functions. Decisions get made on consequentialist grounds. Administrations keep taking the locally reasonable decision to cover up misdeeds. He has made the argument before, drawing on the Charles Armstrong plagiarism case and on the parallel between the three branches of government and the bidirectional character of legal reasoning.
What separates Gelman from most technically prominent statisticians is that he has never mistaken technical authority for moral authority. He does not tell people what to think about politics, policy, or values. He tells them what the data support, and he insists on an accounting of the difference. In an environment that rewards confident, morally inflected, narratively satisfying claims, that insistence subverts. It makes him useful to anyone who wants to know what is known and inconvenient to almost everyone who has built an audience on claiming to know more.
The voice
Gelman writes and talks the way a man thinks out loud at a whiteboard. The voice runs flat and fast. He distrusts grand phrasing and reaches for the small concrete example instead of the sweeping claim.
Start with his diction. He keeps the words short and the syntax loose. He says “this is wrong” and “I don’t buy it” and “I could be making a mistake here.” He avoids the inflated register of the academy. When a technical term shows up, he pauses to deflate it, to say what it means in kitchen English. He prefers the everyday noun to the Latinate one. He would rather say a study doesn’t replicate than dress the same point in the language of robustness failure. The prose feels casual. The casualness hides a lot of control.
His sentences favor the active voice and the present tense, which is part of why he reads as direct even when the argument runs long. He builds by accretion. He states a claim, qualifies it, doubles back, adds an aside in parentheses, then quotes someone at length and answers them line by line. The blog runs on this rhythm. A post often opens mid-thought, as if you walked in on a conversation already going. He uses numbered lists, postscripts, updates stapled to the bottom, and a running cast of recurring examples. The form digresses on purpose. He trusts the reader to follow a tangent and come back.
The rhetoric deflates. His signature move takes a flashy published finding and shrinks it. The beauty-and-sex-ratio paper, the ovulation-and-voting paper, the himmicanes paper, the Wansink food lab. He names the study, names the author, lays out the numbers, and shows why the effect cannot be what the headline says. He coins terms to carry these arguments and the terms stick. The garden of forking paths. The statistical significance filter. Type M and Type S errors. The piranha problem, his argument that a world cannot hold dozens of large independent effects all pushing on the same outcome. The kangaroo, his line about weighing a feather on a bathroom scale while the scale sits on the back of a jumping kangaroo, his way of saying a noisy instrument cannot catch a tiny effect. These phrases do real work. They let him make the same structural point about many papers without sounding like he repeats himself.
He fights, and the affect stays cool. He calls bad work bad. He uses the word fraud when he means fraud and incompetence when he means that, and he draws the line between them. The tone never gets hot. He reports the flaw the way a man reports the weather. The flatness is part of the rhetoric. Heat would invite a fight about manners. Flat delivery keeps the fight on the numbers. He pairs it with steady self-correction. He admits his own past errors, posts retractions of his own claims, says commenters caught him on this. The humility is real and it also functions. It earns him the standing to be hard on others.
Now the spoken man. In talks and on podcasts he sounds like the blog read aloud, only more so. The speech runs fast and associative. He starts an example, interrupts himself to start a second, circles, and lands the point a minute later than a tidier speaker would. He digresses into baseball, into a 1970s study he half-remembers, into something a student said yesterday. He is self-deprecating in a low-key way, quick to say I don’t know and I might be wrong about this. He does not perform authority. He underplays it. His slides, when he uses them, run to rough graphs and screenshots, which fits a man who argues that the picture should carry the argument and the decoration should get out of the way.
He is no stylist of the polished sentence, in speech or print. He is no aphorist working a line until it gleams. The power comes from accumulation and from nerve. A reader who wants a clean essay with a single arc will find him shaggy. The shagginess is the cost of the method. He thinks in public, shows the false starts, and lets you watch him change his mind.
The through-line is a war on false certainty. He hates the move from a noisy result to a confident story. Almost every tic serves that war. The deflating diction, the small examples, the coined terms, the flat tone, the self-correction, the willingness to name names. The manner is the argument. He sounds uncertain about himself and certain about the math, and he wants his audience to learn the same reflex.
The set
Gelman holds court at Columbia and at his blog. The blog runs daily and the comment section serves as the salon. The regulars include Phil Price, Kaiser Fung, Bob Carpenter, Daniel Lakeland, Martha Smith, and a long roster of pseudonymous methodologists. His circle extends to his collaborators, Jennifer Hill, Aleks Jakulin, Aki Vehtari, Michael Betancourt, and to his teacher Donald Rubin (b. 1943), whose causal inference framework supplies much of the technical grammar.
The replication-crisis crowd overlaps almost entirely with Gelman’s. John Ioannidis (b. 1965), Brian Nosek, Uri Simonsohn, Joseph Simmons, Leif Nelson (the three together being Data Colada), Simine Vazire, Daniel Lakens, E.J. Wagenmakers, Anna Dreber, Felix Schönbrodt, James Heathers, Nick Brown, Tim van der Zee, Elisabeth Bik (b. 1966), and Sander Greenland all share the same air. So do philosophers of statistics like Deborah Mayo and statistician-bloggers like Cosma Shalizi, Larry Wasserman, Frank Harrell, Stephen Senn, Christian Robert, Judea Pearl (b. 1936) at the edges, and Kosuke Imai. Nate Silver (b. 1978) sits at the journalistic perimeter. The economists Andrew Eggers, Macartan Humphreys, and Gary King overlap on the causal-inference side.
The set values calibration, transparency, replicability, technical skill at probability and inference, willingness to admit error, slowness of claim, smallness of effect. They want the published record to track the world. They want the standard error to mean what it says. They want preregistration, open data, open code, and they want famous findings checked.
Their hero system rewards the careful auditor. The man who catches the error in the famous paper is the saint. The man who runs the failed replication is the saint. The whistleblower who finds the fraud, Nick Brown on Barbara Fredrickson’s positivity ratio, Brown and Heathers on Wansink, Bik on image duplication, Simonsohn on Lawrence Sanna, is the saint. The patient Bayesian modeler who builds the small honest model that beats the big flashy one is the saint. Gelman canonized this in his forking paths essays and in his insistence that single studies establish almost nothing.
The anti-saints are easy to name. Brian Wansink, Amy Cuddy, Diederik Stapel, Marc Hauser, Daryl Bem, Satoshi Kanazawa, and the more cautious but still suspect Roy Baumeister, John Bargh, and Susan Fiske. Famous studies on power posing, ego depletion, social priming, embodied cognition, and most of behavioral economics from the 2000s sit in the dock. Malcolm Gladwell sits in the dock. TED talks sit in the dock. The New York Times Tuesday science section sits in the dock, though Gelman often appears in it.
Status games run on errors found and frauds caught. The currency is the takedown post, the failed replication, the citation of your blog comment by a journalist, the moment Many Labs posts another null result, the moment Data Colada finds another anomaly. Secondary currencies include the Stan model people fit, the technical paper on multilevel modeling, the textbook that teaches the next generation (Data Analysis Using Regression and Multilevel/Hierarchical Models, Bayesian Data Analysis, Regression and Other Stories), the methodological appendix that solves the puzzle nobody else solved. A man’s reputation rises with each prominent finding he kills.
The set is overwhelmingly male, heavily academic, light on humanities, suspicious of qualitative work, fond of programming, fond of New York and Cambridge and Stanford and Amsterdam and Boston and the Bay Area. The men in it write fast, post often, joke dryly, treat blog comments as a serious form, and treat ad hominem as bad manners while practicing it freely against the named anti-saints. They hold a low opinion of TED, Davos, Aspen, the Edge Foundation, and the celebrity-academic circuit, even as some of them brush against it. They like Tukey, mid-period Fisher, Box, Cox, Rubin. They tolerate but do not love Pearl. They distrust most economists. They distrust most psychologists. They trust each other to find each other’s mistakes and to say so.
The binding glue of the set is a shared confidence that they will be told when they are wrong, by men they respect, and that being told is honor. That is the air they breathe, and that air is rarer than they think.
Their normative claims, stated and assumed, are these. Researchers owe the public honest reporting. P-hacking is a sin. Preregistration is a duty. Failed replications belong in the literature. Effect sizes shrink under scrutiny and that should be expected. Power calculations belong before the study. The burden of proof rises with the surprise of the claim. The press should slow down. Tenure committees should reward rigor over splash. Critics deserve responses. Reviewers should ask for data.
Their essentialist claims are also clear. There is a real distinction between good and bad statistical inference, and trained eyes can tell. There is a real world the data point at, and probability gives partial access to it. Some findings are real and others are not, and the difference is discoverable. The replication crisis describes a real pattern in psychology, medicine, and parts of economics. Bayesian reasoning, done with care, gets closer to truth than the null-hypothesis ritual it replaces, though frequentists in the set, Wasserman, Senn, Mayo at times, push back hard. Method has substance. Numeracy is a virtue and an aptitude, and it can be ranked.
An open question, and how to settle it
Everything above is agreed. What follows is contested, and the contest can be resolved by counting.
Gelman’s credibility rests on symmetry. He says, and his admirers say, that he applies the same standard to work he likes and work he dislikes. The evidence usually offered runs like this. He went after Cuddy, whose work aligned with progressive commitments about female leadership and embodied cognition. He went after the implicit association test literature, which underwrites a large diversity and inclusion apparatus. He has criticized election forecasting done by people who share his politics. He criticizes his own past work. These are not peripheral targets, and a critic optimizing for coalition safety would choose differently.
The counter-case is equally available. His sustained attention clusters in social psychology, behavioral science, nutrition, and management research. It clusters less in economics in its mainstream registers, in biostatistics, and in political science in its liberal registers. His fights over rhetoric, the Vermeule fascism episode among them, do not receive the structural analysis he applies to other people’s rhetoric. A skeptic can assemble a list of fields he leaves alone that looks a good deal like a list of fields where his coalition lives.
Both cases are built the same way. Each side picks the examples that support it. That is the garden of forking paths applied to a career instead of a dataset, and neither side should be believed on this evidence.
The question is answerable. The blog is public, dated, categorized, and searchable back to 2004. Take a defined window, say 2010 through 2025. Code every post that criticizes a named study or a named researcher. Record the field, the journal tier, the political valence of the finding if it has one, and the seniority of the target. Report the distribution against a baseline of what gets published in those fields. That yields a number, and the number can embarrass either side.
Until somebody runs that count, the honest position is that the symmetry claim is unverified in both directions. Everything in the deflationary readings that follows depends on it. If the targeting is symmetric, the coalition-maintenance accounts weaken sharply. If the targeting tracks coalition boundaries, they strengthen. An analysis that asserts the answer without the count is doing what it accuses him of doing.
The deflationary readings
The strongest available reading against Gelman comes from David Pinsof’s Alliance Theory and the family of arguments around it. I take three of them, because a fourth and a fifth restate the first three in new vocabulary.
The first concerns motive. Pinsof’s signaling essay distinguishes offensive from defensive signaling and argues that most signaling is defensive. People are not climbing. They are avoiding a fall. Read that way, Gelman’s policing looks different from what a coalition-maintenance account predicts. The fear of being the man who let junk science slide, who stayed quiet while Wansink ran his food lab, who said nothing while power posing reached millions, fits his behavior better than an ambition to be recognized as the most rigorous statistician of his generation.
The texture supports it. He does not build a brand. He seems unable to stay quiet when he sees something wrong. The blog’s tone compulsive. He posts corrections, qualifications, responses to responses, follow-ups on follow-ups. The volume and the consistency look like a man who cannot stop noticing errors. A pure offensive signaler would select more carefully. He would pick targets that maximize visibility and minimize coalition cost.
The defensive frame also reads the Cuddy fight better. The New York Times Magazine cast his criticism as bullying and as pleasure taken in the destruction of a career. His answer was that the accuracy of published claims is not private. That is a defensive move: an attempt to avoid being the man who knew a paper was weak and said nothing, who took part in the private correction system that keeps errors circulating while protecting reputations. The shame he defends against is the shame of complicity.
This matters for prediction. A strategic signaler adjusts when the rewards change. A man running on defensive motivation keeps going regardless, because what he defends against does not go away when the status game shifts. That is a falsifiable claim about his next decade, and it is the most useful thing this family of frameworks produces about him.
The second concerns timing. Pinsof’s essay on status traces how status games collapse and re-emerge in inverted form. When common knowledge sets in that a game is a game, the hierarchy flips. Players who accumulated rank through the old signals start to look conniving. People who did the opposite start to look humble.
The pre-crisis social psychology game rewarded impressive significant results from small samples, emotionally resonant findings, policy-relevant claims, and TED-ready narratives. Gelman had been doing the opposite. He reported uncertainty when the market paid for confidence. He demanded replication when the market paid for novelty. He published wide intervals when the market paid for clean results. He was losing. The crisis collapsed the game he was losing, and his accumulated record of doing the opposite turned into the most valuable credential in the wreckage.
He did not accumulate it in anticipation. He accumulated it because he could not do otherwise. The collapse rewarded it as though he had planned it, which is what the theory predicts: the man who profits most from a collapse is often the one least invested in the game that collapsed.
The theory then makes an uncomfortable forward prediction. The anti-game he won is now a game, played in the dark by a community that has internalized pre-registration, open data, and honest uncertainty as sacred values. Signs of the instability are already visible. Pre-registration is widespread enough that strategic pre-registration has appeared. Open data is widespread enough that data gets shared in formats that satisfy the letter and block replication. Fluency in the vocabulary of rigor now works as a coalition signal. Gelman has noticed some of this and written about it. He sees the junk science game clearly because he is outside it. He sees the rigor game less clearly because he is inside it and winning.
The third concerns the theory of the problem. Pinsof’s misunderstanding essay argues that intellectuals mislocate the causes of human problems, because the misunderstanding story makes intellectuals indispensable. If researchers produce bad work because they do not understand statistics, the man who teaches better statistics performs an essential service. If they produce bad work because the system rewards it, statistical education treats a symptom while the disease operates where the educator cannot reach.
Applied to Gelman, this cuts his output in half along a seam he does not draw himself.
One half operates on the misunderstanding model. The blog, the textbooks, the public criticism of individual papers all assume that if enough researchers grasp Type S and Type M errors and the garden of forking paths, the field improves. The other half operates on incentives. Pre-registration removes the payoff to exploring until significance appears, because the exploration goes on record. Open data removes the payoff to hiding inconsistencies. Registered reports remove the payoff to running many studies and reporting the lucky one, because the journal commits before results exist. None of these teaches anybody anything. Each changes what pays.
The prediction is that the second half does the durable work and the first half does less than it appears to. The evidence sits in Gelman’s own August reply. Inside the credentialed system, where the incentives changed, reform happened. Outside it, where they did not, the problem got worse. Huberman has every reason to produce confident health claims and none to produce accurate ones, because his audience cannot evaluate accuracy and pays for confidence. His audience never agreed to play the specificity game. Gelman’s demand for precision arrives there as an uninvited disruption of something working as intended, and he has no lever to pull.
There is a partial defense available, and it is worth stating because it complicates the split. The blog functions as a reputation tax. Knowing that Gelman might write about your paper creates a small real cost. The public criticism of Wansink and Cuddy was never aimed at those two. It was aimed at everyone watching who could see that public methodological failure now carries a price. On that reading the educational work is incentive work wearing a teacher’s clothes, and the seam is narrower than the theory suggests.
The hardest version of the deflation turns on Gelman himself. The intellectual who derives standing from naming other people’s errors has an interest in finding errors worth naming, and that interest runs independent of whether naming them improves science. His profile rose because the crisis gave his methodological work a public stage. None of that makes the critiques wrong. They were right. Being right and being positioned to profit from being right are compatible.
Two more frames deserve short notice.
Stephen Turner’s convenient beliefs framework asks which beliefs a man’s position makes comfortable. For a methodologist the most comfortable belief is that the crisis is methodological. If the crisis is methodological, the person with the best methods is the most important person in the room. If it is structural, if the publication system and the career incentives and the prestige economy select for unreliable findings regardless of anyone’s statistical sophistication, then better methods are necessary and insufficient, and the methodologist becomes one voice among many in a conversation about institutional design where statisticians hold no special authority.
Gelman knows this. He has written about incentive structures, publication bias, and the sociology of science. The framework does not accuse him of hiding it. It predicts that the pedagogical emphasis persists in his primary output because that emphasis preserves his function, and that prediction is checkable against the same corpus the targeting count would use.
Turner’s framework produces a second observation about the blog. Gelman describes it as open intellectual exchange, which it is. It is also a platform no journal can match for speed and reach. It lets him set the terms of debate, decide which papers deserve scrutiny, and determine which errors get public attention. He is the editor of a shadow journal with no editorial board and no accountability beyond his own judgment. That his judgment is usually good does not touch the structural point.
Jeffrey Alexander’s work on cultural trauma and on Watergate as democratic ritual supplies the last useful frame. Alexander argues that traumas are constructed by carrier groups doing representational work: naming the injury, identifying the victims, connecting victims to a wider audience, assigning responsibility. He names the scientific arena as one site where this happens, which is easy to miss because the scientific arena presents itself as the place where traumas get diagnosed.
The replication crisis fits the description. A set of events, failed replications, questionable practices, a few frauds, became a civic crisis for academic social science through the work of a carrier group that included Gelman, Nosek, Simonsohn, Vazire, and a network of methodologists. Before their work the failures were scattered anomalies inside subfields. After it they were symptoms of a general pathology threatening the legitimacy of social science.
Alexander’s framework does not call this manipulation. It calls it what carrier groups do when they work. The construction produced real change. It also produced what successful constructions produce over time: sacred values that outlive the argument and get implemented by people who relate to them through institutional pressure. Pre-registration became a virtue marker. Open data became an honesty marker. Power analysis became a requirement. Each has generated its compliance ritual. Gelman has pushed back against the drift from reform to ritual. The framework predicts he cannot stop it, because the drift is how constructions become institutions.
Alexander also explains something about the blog’s authority that the arguments alone do not. The Senate Watergate hearings created a space where ordinary political rules were suspended and statements carried weight they would not carry in mundane politics. The blog does something similar inside academic statistics. The consistent voice, the daily rhythm, the commenter community with its own conventions, the shared vocabulary that marks insiders, the cross-references reaching back twenty years: these give a post more force than its argument alone would command. A paper gets classified as an instance of a known pathology and the classification sticks. A researcher becomes associated with polluted practice and the association follows him. Whether the markings are accurate is a separate question the framework does not settle. What it settles is the character of the authority. The blog amplifies its arguments without necessarily distorting them, and the amplification is structural.
What the deflation cannot do
Now the objection. Ernest Becker (1924-1974) named, in The Denial of Death, the work a man’s creed performs, the holding-off of death through service to something that outlasts him. Gelman’s something is the self-correcting record, the slow public machine by which inquiry catches its own errors and grinds toward truth across generations. His methods, his students, the norms he pressed on a field, these go on after him, and the going-on is his answer to the grave. The story his life tells is that science is fragile, that the crisis threatened to rot it, and that his criticism defended something larger than any career.
A reductive reader says the story is a cover, that under the talk of integrity sits the ordinary fear of slipping down the ladder. The reductive reader has not earned the claim. Becker’s point was never that the immortality project masks a baser motive. The project is how the motive lives in an animal that knows it will die. The hunger for significance and the love of the enterprise are the same hunger, and to call the nobler name a disguise is to claim a knowledge of another man’s heart that no evidence supplies. That is the one move Gelman spent his life teaching us to distrust.
Sit with that. To deflate him, to announce that his integrity is status anxiety dressed for church, is to do the thing his whole career condemns. It is a finding with no power behind it, a confident story reverse-engineered from a man’s success, the garden of forking paths run on a biography instead of a dataset. You can always find the path that makes the honest man look like a careerist, the way you can always find the subgroup that turns the null result significant.
He taught the field to ask what we would believe if the study had come out the other way. Ask it here. Had the disintermediation never come, had his kind of science kept its grip on public belief, nobody would read his integrity as a cover for status fear, because there would be no falling status to explain it by. The deflation depends on the outcome it pretends to diagnose. By his own time-reversal test it fails.
So here is what survives and what does not. The three readings above survive as descriptions of function, not as claims about motive. Defensive signaling describes what his behavior accomplishes and predicts what he will do next. Status game collapse describes why his standing rose when it did. The incentive analysis describes which half of his output changes behavior. None of the three requires that he be insincere, and none of them supplies the evidence that would license saying so. What does not survive is the further move, the one that treats sincerity as the most effective concealment and thereby makes the case unfalsifiable.
He won the war he fought. Inside the academy the reforms took: the pre-registration, the open data, the death of the lonely underpowered study waved through on a lucky p-value. The victory arrived as the ground gave way beneath it. The bridge from rigorous research to public belief, the science journalism and the popularizers and the lectures that carried findings from the lab to the living room, gives way, and into the gap pour the direct-to-audience health influencers who need no credential and answer to no review, whose authority is reach and warmth and the parasocial trust of millions. Gelman perfected the instrument and the concert hall emptied. He is right inside a house whose writ no longer runs where most people form their beliefs. His reply names this without flinching, and a lesser man would have told himself a happier story.
A quieter cost sits beneath that one. The discipline that forbids overclaiming forbids the verdict too, the meaning, the thing a frightened public wants. A man deciding how to live, whether to fear the diagnosis or take the supplement or trust the shot, comes to Gelman and receives a probability interval and a warning that the study was underpowered, which is the truth and is not the bread he came for. The influencer hands him certainty and a plan. Gelman hands him the honest width of the unknown. The honest width is worth more and feels like less, and in a market for feeling, the man who sells the truth about uncertainty is selling the one thing a frightened animal is built to refuse.
Which returns us to the April email. Asked about the unwritten rules of his own campus, he answered that he knows of none. The answer was sincere. He is a senior tenured professor, aligned with the dominant coalition of his institution, credentialed in ways that lift him above most departmental conflict, and prominent enough that his standing absorbs the low-level social penalties that enforce tacit norms on graduate students, junior faculty, and undergraduates. Those rules are not experienced uniformly. They press hardest on people whose position is precarious, whose coalition membership is uncertain, or whose views sit near the boundary of what the dominant formation tolerates. He sits nowhere near that boundary. Of course he has not felt them.
This is the finding. Michael Polanyi called it tacit knowledge, and Gelman knows the concept well enough to have written about how statistical judgment resists reduction to explicit rules. He can apply it everywhere except to the one place where the application would cost him something, which is his own formation. The gynecologist who hears the patient report and files it as confounded is operating inside a framework that filters what reaches him, and the framework could not work if he could see it working.
The others in this gallery have a blind spot they cannot find. Gelman is the strange case who sees most of the board, including the square his own king stands on. He runs the skepticism on himself, corrects his own old work, and names the obsolescence creeping toward his method without dressing it as another man’s fault. The cut is that seeing does not save him. Rigor cannot manufacture the public trust that rigor once earned, and the virtue that built the bridge holds no tool to rebuild it after the culture stops prizing the virtue. He can describe the washing-out of the road with accuracy. He cannot pave it with description.
So the figure stands. The honest accountant of what can be known, the man who made restraint heroic in a field that pays for confident lies, and who turned his skepticism on himself when the others turned theirs only outward. His hero is the un-self-deceived inquirer. His immortality is the self-correcting record. He is doing the most honest work in the building. The building empties. He keeps the books straight anyway, which is either the last virtue or the first one, and is in any case the only one he was ever willing to claim.
