Cohen Passes His Own Audit

Having recounted Philip N. Cohen’s audits of the American Sociological Review and found them overstated, the obvious next question is whether he meets the standard he applies to other people.

He does.

The frame is his curriculum vitae. Screening it for peer-reviewed empirical articles with quantitative analysis published between 2015 and 2025 gives twenty papers. That window covers the years in which open science became half his public program.

The rule is the one frozen for the ASR recount. Count an article when publicly accessible code appears, without executing it, to cover the article’s quantitative results, and the data are either included or obtainable by outside researchers through a documented procedure that does not run through the authors. Restricted data do not disqualify; discretionary data do. Partial coverage does not count. No code was run.

The first result was unexpected. Not one of the twenty articles appeared in a journal that required public deposit at the time of publication.

That took some establishing. Sociological Science requires a replication package now, and its policy applies to manuscripts submitted after April 1, 2023; Cohen’s article there appeared in January 2016. Wiley announced its tiered framework in September 2017, after his 2016 article in The Sociological Quarterly, and left the tier to each journal. Elsevier encouraged rather than mandated. Springer Nature says its general policy requires a data availability statement and introduces no sharing mandate. Sage operates tiers and distinguishes journals that require public sharing from journals that encourage it; no Sage title in this set was in the mandatory group at the relevant date.

Seven of the twenty were subject to a statement requirement.

So every package Cohen posted across eleven years was posted without a journal compelling it. Thirteen of the twenty articles have one. Sixty-five percent.

Under the same rule, the 2025 volume of ASR comes in at forty-four percent. Cohen’s own record beats the journal he has been auditing.

On sole-authored papers he is seven for seven. Every one. He is also one for one as first author of a multi-author paper. The gaps sit further down the author list: four of nine as second author, one of two as third, none of the one where he is fourth.

That distribution is what you would expect if deposit tracks whoever ran the analysis and assembled the files. It also means the seven-for-seven figure is the one that speaks to his own practice, and the rest speaks to his coauthors’.

The trajectory tracks his public program. Through 2017 he is one of five. From 2018 forward, the years in which he launched SocArXiv and began arguing in public about transparency, he is twelve of fifteen. His two 2015 papers, with Jeehye Kang in the Journal of Family Issues and with Lucia Lykke in Social Currents, have no packages and no data availability statements. His 2025 article on views of falling birth rates carries a statement pointing at the Pew source data and an OSF repository for the code.

The man changed his practice when he started telling other people to change theirs. That is the least he could have done and it is more than the record shows for most people who argue about this in public.

Three cautions. Twenty articles is a small denominator and the authorship cells run to one and two papers; nothing here supports a claim about how coauthorship affects deposit in general. The venue finding is a negative search across twenty journal-years, well supported by the publisher framework documents and still an absence rather than a positive record. And the rule I applied is his 2023 rule, which is stricter than his 2025 one; under the looser standard his rate would be higher.

The easy charge against a man who audits other people is that he would not survive his own audit. Cohen survives it. On the papers where he alone decided, he survives it completely, in years when no journal asked him to.

That makes the criticism in the other piece more interesting rather than less. The problem with his 2025 audit is that he changed the ruler between two measurements and read the difference as movement, which is a failure of the method rather than of the man, and which the man’s own thirty years of work would have caught in anybody else.

The ASR case ledger gives the coding rule this census applies.

What follows is the venue ledger behind the census: for each of the twenty articles, the journal’s requirement at the date of publication and the source establishing it. All checks August 26, 2026.

Jeehye Kang and Philip N. Cohen published in the Journal of Family Issues in September 2015. No public deposit, code deposit or pre-publication check requirement was located for that date, and the article carries no data-availability statement. Sage’s later research-data framework distinguishes journals that require public sharing from journals that encourage it, and no evidence places this title in the mandatory group in 2015.

Lucia C. Lykke and Philip N. Cohen published in Social Currents in September 2015. No deposit, code, statement or check requirement was located, and the article carries no data-availability statement.

Cohen published alone in Sociological Science on January 25, 2016. The journal has no deposit requirement applicable to that date; its mandatory replication-package policy applies to manuscripts submitted after April 1, 2023.

Jill E. Yavorsky, Philip N. Cohen and Yue Qian published in The Sociological Quarterly, online July 20, 2016. No public-deposit requirement was located. Wiley announced its tiered framework on September 14, 2017, fourteen months later, and left each journal to choose among encouragement, expectation and mandate.

Jeehye Kang and Philip N. Cohen published in Social Science and Medicine in April 2017. No public-deposit requirement was located. Elsevier’s November 30, 2016 initiative implemented data-citation standards to encourage sharing without mandating deposit, and its general policy describes sharing as encouraged unless a journal adopts a stronger rule.

Philip N. Cohen and Joanna R. Pepin published “Unequal Marriage Markets” in Socius in August 2018, with a voluntary OSF package. No public-deposit requirement was located. The contemporaneous Socius article “Institutionalizing Transparency” discusses stronger journal policies as reforms to be adopted rather than describing a rule the journal already had.

Cohen published “The Coming Divorce Decline” in Socius in August 2019. No public-deposit requirement was located; Sage’s guidance as catalogued in September 2019 supported data sharing rather than requiring it.

Brandon G. Wagner, Kate H. Choi and Philip N. Cohen published in Socius in December 2020 and supplied an OSF package. No public-deposit requirement was located.

Joanna R. Pepin and Philip N. Cohen published “Nation-Level Gender Inequality and Couples’ Income Arrangements” in the Journal of Family and Economic Issues in 2020. The title sat under Springer Nature’s general policy, which asks authors to supply documentation on request rather than to deposit publicly.

Cohen published alone in the European Journal of Environment and Public Health in June 2020 and linked data and code at OSF. No deposit, code or check requirement was located.

Jeehye Kang, Philip N. Cohen and Feinian Chen published in Child Indicators Research, online in 2020. No mandatory public-deposit rule was located. Springer Nature’s single statement policy arrived on March 22, 2023, more than two years later, and requires statements rather than deposit.

Cohen published “The Rise of One-Person Households” in Socius in December 2021, with data and code at OSF. No public-deposit requirement was located.

Cohen published alone in Social Determinants of Health in February 2022. No deposit, code, statement or check requirement was located. The OSF package was already public with the December 2021 preprint.

Cohen published “Rethinking Marriage Metabolism” in Population Research and Policy Review in 2023. A data-availability statement was required or being implemented under the March 2023 policy, which encourages repository deposit while permitting statements that describe restricted data. Public deposit was not required. Package.

Mónica L. Caudillo, Andrés Villarreal and Philip N. Cohen published in The ANNALS of the American Academy of Political and Social Science in 2023. No mandatory public-deposit rule was located for this title at this date, and the published record supplies no qualifying package.

Joanna R. Pepin and Philip N. Cohen published in Socius in 2024. A data-availability statement appears in the article and the replication materials are linked voluntarily at GitHub. No deposit mandate was verified.

Hao-Chun Cheng and Philip N. Cohen published in Asian Population Studies in 2024. Taylor and Francis’s basic policy encourages a statement and data sharing without requiring deposit, and no evidence places this title in a mandatory tier. Package

Ansgar Hudde and Philip N. Cohen published in Socius in 2024. The statement links code, output, GSS and IPUMS data, with materials at OSF. No mandatory public-deposit rule was located.

Hangqing Ruan, Melody Ge Gao, Yajie Xiong and Philip N. Cohen published in Socius in 2024. A statement rule may have applied; no mandatory deposit rule was located, and no qualifying replication package was located as of the check date.

Cohen published alone in Socius in 2025. The statement links the Pew source data and code at OSF. Sage’s policy prescribes what a statement must contain and distinguishes that from the stronger rule imposed by journals requiring repository deposit.

Posted in Sociology | Comments Off on Cohen Passes His Own Audit

Andrew Gelman & the E-Personality

In June 2007, three years into Statistical Modeling, Causal Inference, and Social Science, Andrew Gelman (b. 1965) wrote a short post about blog personalities. He and Seth Roberts (1953-2014), who interviewed him about blogging that same spring, were both aggressive conversationalists face to face, he said, and both of them turned into someone else online. His own online self was the mellow one. He gave a reason. An academic with a reputation has more to lose than to gain from gratuitous controversy, so when he criticized something he could substantiate, he did it in a mild voice.

Elias Aboujaoude (b. 1971) published Virtually You: The Dangerous Powers of the E-Personality four years later, in 2011. His argument is that sustained life online produces a second self, more assertive and less restrained than the offline person, and that the second self does not stay online. He gives the traits chapters: delusions of grandeur, narcissism, ordinary everyday viciousness, impulsivity, infantile regression, the illusion of knowledge, the end of privacy. The closing chapter names the migration virtualism, the process by which the offline man drifts toward the online one.

Gelman inverts the opening premise. Every other subject I have run through this framework, such as Paul Bloom, Philip N. Cohen, Amy Wax and Nathan Cofnas, gets louder online. Gelman got quieter on purpose, and then his influence boomed.

The prudential reason he gave is incomplete, and his own archive shows it. Seventeen years before the blog personalities post, in a Harvard dissertation on image reconstruction for emission tomography, he wrote in the same register. The PET study behind the thesis had five subjects, and he says in the text that the sample was too small to estimate the effects of interest. He reports in his own conclusion that an implementation he does not present indicates the standard model of PET data is incomplete in practice. He ran it, it failed, and he told a reader who was never going to check. There was no reputation at stake in 1990 and no audience to manage. The mild voice was not a policy adopted for the Internet. It was already how the man wrote when nobody was reading, which means the blog did not create the restraint. It found it and put it to work.

Consider first what Aboujaoude calls grandiosity — the swelling of what a man believes he can do and how far his word should carry. In a university, Gelman is a statistician and a political scientist. Those are his departments and his training and the rooms where his authority is chartered. On the blog he passes judgment on psychology, epidemiology, nutrition, public health, education, medicine, polling, economics, and whatever paper landed in his inbox that morning. Noah Smith put the result to him in a 2022 conversation: the frightening sentence in academic life is that Andrew Gelman just blogged about your paper.

An expansion of jurisdiction on that scale has no offline equivalent. No committee appointed him. No editor commissioned him. No discipline credentialed him to referee the others. The blog gave a man trained in one thing standing to rule on twenty, and the ruling sticks because the technical objection travels across fields even when the substance does not.

That is where the illusion of knowledge enters, and where Gelman is at once the phenomenon and the argument against it. Aboujaoude’s worry is that fast access to information gets mistaken for mastery, and that the network trains us to know less and less about more and more.

A post from December 2015 shows the shape. A reader sent him a biology paper and asked whether he believed it. Gelman’s answer amounted to this: seventy-one subjects split four ways, no. He said in the same breath that he did not know the field and could not evaluate the biology. He ruled anyway, from the structure of the evidence rather than the content of the claim.

Read one way, that is the Internet expert Aboujaoude warns about, pronouncing on a literature he has not read. Read another way, it is the whole point. A design that cannot support an inference cannot support it in any field, and a man who understands designs can say so.

Gelman consistently marks the edge of what he knows and keeps marking it. I am no expert here. I do not know. Here is what bothers me. In January 2011, asked how to teach introductory statistics better, he said he did not know how, and added that one of the pleasures of tenure and blogging was being free to say so rather than bury the uncertainty for publication. The medium that Aboujaoude says inflates confidence is the medium Gelman uses to publish his uncertainty, because the journal will not take it and the seminar will not reward it.

So the illusion of knowledge cuts twice. The blog enlarges the territory across which he can pronounce, and his central intellectual commitment is the deflation of pronouncements.

The viciousness chapter fits with less resistance. Aboujaoude’s argument there is that anonymity is not required for disinhibition. Named people harden online too, because the ordinary brakes, the face across the table, the warmth, the manners, all go missing.

The mature blog runs on a vocabulary of ridicule. Power pose. Himmicanes. Gremlins. The coinages are funny and they are meant to be. They turn a methodological failure into a comic object with a handle, and the handle is why the lesson travels. Gelman has defended the practice on those grounds. During the 2016 fight over methodological terrorism he said that a little mockery makes a point memorable, and he has admitted writing silly posts because he wanted to entertain. The New York Times, describing the blog, noted that his critiques get read with a certain pleasure at other people’s expense.

He has set the reasoning out at length, and it holds up on its own terms. Ridicule is vivid. Vivid teaches. Status protects bad work, and mockery punctures status.

Would the coinage arrive if the target were sitting three feet away? Sometimes, probably. Always, no. The screen changes what each move costs. The target’s embarrassment is a name on a page. The audience’s amusement arrives within the hour, in the comments, in the reposts, in the phrase entering circulation. What reads as cruelty across a seminar table reads as wit when the man is a hyperlink.

The coinages are the mockery, and they attach to particular targets. Count them. If the ridicule vocabulary clusters on the famous and the tenured and the TED-adjacent, then the disinhibition runs along a status gradient rather than a random one, and the finding is about who the form licenses him to mock rather than about whether he is mean. That count belongs with the larger audit of his targeting, and it would settle a question that this essay can only raise.

Narcissism needs care, because the word carries a clinical meaning that has no business here. Twelve thousand posts and two hundred thousand comments, written mostly by one man, construct a world where that man’s attention is the organizing principle. A paper enters because he found it interesting. A dispute continues because he kept thinking about it. A researcher recurs for a decade because the case became useful in his vocabulary. Wansink is not a subject on that blog. Wansink is a category.

Gelman publishes corrections. He links to people who say he is wrong. His comment threads are full of readers telling him so, and he answers them, and he changes his mind in public, and he says which commenter caught him. In 2019 he credited part of his methodological development to answering questions from strangers on the Internet. An e-personality organized around admiration does not run an open thread for twenty-two years and let it correct him in front of the audience.

Aboujaoude’s Internet rewards sending before thinking, action taken before the information is in. Everything about the blog’s surface says impulse. The prose is colloquial. A post opens mid-thought, as if you walked in on a conversation. Somebody just sent him this. This seems ridiculous to me. Thought appears to become publication with nothing in between.

Then look at the machinery. He was already running a backlog in 2007. By 2015 he was posting a month or two in arrears and arguing that the lag improved the result, because by the time a post ran he had seen more and could revisit it. In April 2018 he mentioned that an item written in September had waited until then for a slot. On August 15, 2026, he announced that the queue had reached a full year. Posts written now appear in August 2027. He still bumps timely material forward, and his co-bloggers publish when they like.

A man who has turned the instant medium into a periodical is not the man the impulsivity chapter describes. Something more interesting is going on. The voice still says I just saw this, and the publication system says nothing of the kind. Spontaneity has become a style rather than a condition. The prose performs immediacy that the calendar has removed, and the performance is so consistent that most readers have no idea they are reading a man’s thoughts from last summer.

Regression comes out mixed. Aboujaoude means the childlike register the browser invites, the emoticon standing in for the sentence, and, more seriously, the way the short forms excuse a writer from building an argument that has to hold together across pages.

The playful half applies. The joking titles, the running gags, the feuds, the recurring parables all let a chaired professor at Columbia behave in ways The Annals of Statistics would not permit. The blog gives him room to play, and he takes it.

The serious half fails. Posts run to thousands of words. Arguments continue through hundreds of comments. A post from 2011 refers back to one from 2006 and forward to one from 2019. Papers have come out of comment threads. Readers arrive with counterarguments and he revises. The retro feed he set up in 2019, which reposts the archive from the beginning at one entry every eight hours and will take a decade to catch up, is a monument to how much of it is worth rereading. The Internet fragmented most intellectual life. He used it to rebuild a seminar that never adjourns.

Which brings the argument to virtualism.

Aboujaoude’s last move is that the online self migrates. The offline man becomes, over years, an approximation of the man he plays on the screen. In most cases that means a person gets harsher, faster, surer, and does not notice.

In Gelman’s case the migration went somewhere stranger.

In 2007 he had an alter ego who was calmer than he was, adopted for prudential reasons he could state out loud. That alter ego published every day for two decades. It acquired a vocabulary, a cast of recurring characters, a body of parables, a readership that arrives already knowing the references, and a reputation that arrives at a researcher’s inbox before the post does. Somewhere in that accumulation the alter ego stopped being a manner and became a position. The mellow blogger turned into the man Noah Smith could call the scourge of bad statistical papers, the man whose attention functions as a sanction.

Andrew Gelman working through journals, seminars, and referee reports can tell three people that a paper is bad. Andrew Gelman on the blog can tell thousands, without an editor, without a referee, without a discipline’s permission, on a schedule he sets, in a voice that admits doubt, with a comment section that lets the accused answer. That is a form of standing no university confers and no journal can revoke.

The public Andrew Gelman is by now a character, and the character is more stable than any man. He hates unearned certainty. He pounces on noisy studies. He loves multilevel models and hates bad graphs. He distrusts prestige. He revisits old enemies. He says I don’t know. He tells you when he was wrong. Twenty-two years of posts, thousands of commenters, and a permanent archive built him jointly, and no single author, Gelman included, can now revise him. The archive does not forget.

Aboujaoude expects the medium to enlarge the man’s aggression. Here it enlarged his jurisdiction. The 2007 inhibition was real and it held for years, and it turned out to be the thing that made the expansion possible, because a mild voice applied to a hard claim for two decades earns a standing that a loud voice never would. He was careful because he had a reputation to protect. The care built a larger reputation, which bought a larger license, which he now exercises across every field that publishes a number.

There is a caution in that for anybody who admires the result. The office he built is available to him and to almost nobody else. Sustained public criticism of famous work requires emotional stability, financial independence, tenure, technical credibility, a second discipline to stand in, textbooks everyone assigns, and no need for the target’s approval. A graduate student who wrote the same posts would likely be unemployable.

Aboujaoude worries that the archive holds us to selves we have outgrown. Gelman’s archive does exactly that, and he treats it as an asset. He can point at what he said in 2004 and show that the equipment predates the crisis that made it valuable. Most men would find a twenty-two year record of their opinions a liability. A man whose whole method is to check what a claim looked like before the result came in finds it the best evidence he has.

Posted in Andrew Gelman | Comments Off on Andrew Gelman & the E-Personality

‘(1) “Do you think the culture of research has genuinely changed since the replication crisis became widely discussed, or has it mostly generated new compliance rituals around pre-registration and open data while leaving the underlying incentive structure intact?, (2) Regarding Columbia University, “is there a statistical or social scientific way of understanding how institutions lose the ability to accurately perceive their own situation?”’

Andrew Gelman writes:

My reply:

1. I’m loath to give an answer about the changes in the culture of research because I have not studied this systematically. My impression is that, yes, there’s more skepticism and less acceptance of noisy N=38 papers in psychology, etc., and less toleration for unfalsifiable evolutionary psychology and that sort of thing. On the other hand, perhaps this has just shifted from the science establishment to social media. Ten or fifteen years ago, there was a pipeline (partly abetted by Jeffrey Epstein) from researchers at top universities to publication in top journals to books, NPR, Ted, Gladwell, Freakonomics, etc., and lucrative speaking and consulting gigs. So you get people like Marc Hauser or Albert-Laszlo Barabasi or Brian Wansink or Dan Ariely doing the basic research (such as it is), academic middlemen such as Steven Levitt and Cass Sunstein as promulgators, and the universities, journals, and prestige news media as part of this system (as for example here: https://statmodeling.stat.columbia.edu/2023/08/31/the-variation-ignoring-junk-science-thats-promoted-by-association-for-psychological-science-and-related-academic-celebrities-its-like-a-poker-player-thinking-okay-if-push-all/).

Nowadays, though, social media runs on its own steam, and the models for academic junk science are researchers such as Andrew Huberman and Dr. Oz, who cut out the middleman and promote junk science directly, sell supplements, etc. And social media is full of fake news and AI slop. They don’t really need NPR, Ted, Gladwell, Freakonomics, etc., anymore; they can do it on their own. So, in short, yes, I do have the impression that science has reformed from the bad old days of 2010-2015 (about which, see this article with Simine Vazire: https://sites.stat.columbia.edu/gelman/research/published/jmmss-3062-gelman.pdf), but maybe the public intellectuals don’t need academic science anymore; they can just make up whatever they want on their own.

2. My take on Columbia is similar to my take on many institutions, which is that they have an executive function but minimal legislative or judicial functions; I discussed this here: https://statmodeling.stat.columbia.edu/2018/01/19/lesson-charles-armstrong-plagiarism-scandal-separation-judicial-executive-functions/ and here: https://statmodeling.stat.columbia.edu/2025/11/11/from-the-three-branches-of-government-to-the-bidirectional-nature-of-legal-reasoning-in-a-way-that-is-similar-to-how-statistics-works-and-should-work-in-the-real-world/. As a result, their decisions are made on consequentialist rather than proceduralist gounds, and over and over again the administration takes the seemingly reasonable decision to cover up misdeeds.

Ford posted this discussion on his blog. It was kinda weird seeing myself discussed as a sociological object, but, fair enough, I’m a public figure, and people can say what they want as long as they don’t misrepresent my writings or claim that I said something I never said.

Ford’s assessment is accurate that I’m not very good at strategic behavior so often I don’t even try. It’s similar to how I’m a bad negotiator so usually I’ll just try to make my goals clear and not try to optimize, following the “Getting to Yes” principle that the main thing getting in the way of smooth negotiation is ignorance of other people’s goals. I think back to various successful and botched negotiations I’ve been involved with in the past, and almost always the problems come with struggles over details without there being clarity on the goals of the different parties.

Reform worked inside the academy, and the demand for junk science routed around it. The old pipeline needed universities as collateral: a lab, a journal, a trade book, NPR, the speaking fee. Raise the cost at the journal step and the trade moves to channels with no gatekeeper at all. Andrew Huberman (b. 1975) and Mehmet Oz (b. 1960) sell direct. That is a displacement argument, and it is testable in principle. You could count how many of the top-selling behavioral-science trade books of 2010 to 2015 rested on a named peer-reviewed finding, and compare with the top health and self-improvement podcasts of the past three years.

The weakness in his own story is the brand. Huberman trades on Stanford. Oz traded on Columbia. The middleman got cut, the imprimatur did not. If the university name still does the underwriting, then academic reform is doing less work than the reform narrative suggests, because the asset being rented is the affiliation and the affiliation survives the correction of the paper.

I asked whether the incentive structure changed. Gelman answered about the culture of reception and the fate of the promoters. Tenure, hiring, and grant review still pay for volume and novelty. Pre-registration is cheap to satisfy and easy to hollow out. So my dichotomy stands mostly unaddressed.

The Columbia answer is the richer one and the least developed. His model: institutions have an executive function and thin legislative and judicial functions, so decisions get made on consequences and covering up keeps looking like the reasonable local choice. That explains chronic self-misperception without needing anyone to be stupid or wicked. An institution with no internal adversarial process produces no findings of fact. It has no record of what it did that was not written by the people who did it. Its model of itself is assembled from its own communications office. Every step of the cover-up is defensible, and the aggregate is an organization that cannot see itself.

He skipped the statistical half of my question. There are answers available in his own idiom. Selection: the administration hears from the channels that select for agreement, and treats the resulting sample as the population. Measurement: the quantities that get measured are the ones the executive already controls, so the dashboard is a mirror. Forking paths applied to self-assessment: with enough readings of the same events, some reading always vindicates the decision that was already made.

He confirms my characterization of him on the record. He says he is bad at strategic behavior, does not try, and negotiates by stating his goals rather than optimizing. A subject ratifying a profile’s central claim in his own words is rare. His line about being discussed as a sociological object, and then granting the fairness of it, is the same trait operating in real time.

Posted in Andrew Gelman | Comments Off on ‘(1) “Do you think the culture of research has genuinely changed since the replication crisis became widely discussed, or has it mostly generated new compliance rituals around pre-registration and open data while leaving the underlying incentive structure intact?, (2) Regarding Columbia University, “is there a statistical or social scientific way of understanding how institutions lose the ability to accurately perceive their own situation?”’

What a Public Audit Does (V4)

In November 2020 Philip N. Cohen posted an inspection of the American Sociological Review. He went through the quantitative research articles in four issues of Volume 85, looked for links to replication materials, found that four of fifteen provided them, and named the articles that did not.

In the same post he set out a theory of why researchers share. Authors share when they are personally compelled, when they are normatively compelled, or when they are formally compelled. The audit was an exercise in the first two. It made a public record of who had shared and who had not, on the premise that a scholar who sees his own paper listed under nothing will behave differently next time, and that a discipline that sees the list will develop an expectation.

He repeated the exercise in January 2024 and again in December 2025.

This note asks what those audits produced. It reports two tests of the personal-compulsion route, both returning zero events in every arm. It reports a recount of the 2025 audit under the 2023 standard. It reports the journal’s policy history across the period, which rules out formal compulsion. And it reports a detection problem, found by accident, that bears on every audit of this kind including this one.

None of these is a large finding. Together they narrow what a public audit can be doing.

The three audits do not use one rule. In January 2024 Cohen sorted twenty-five quantitative articles into three groups. Eight had what he called a full package, meaning data and code sufficient to replicate the paper, judged by inspection rather than execution. A further group supplied something less. The rest supplied nothing. His table and his prose disagree on the middle group, giving five in one place and six in the other, with the residual moving accordingly.

In December 2025 the middle group disappears. The positive category now holds articles providing packages of code and data, or information about getting the data. He counts materials hosted on personal websites and excludes materials described as available on request. Twelve of eighteen articles clear that bar.

He states the 2025 rule in the post. He then compares the resulting figure against the 2023 figure and reports that the journal is making progress.

I (meaning Claude and ChatGPT) recoded the eighteen 2025 articles under the 2023 standard and recoded his eight 2023 positives under the same rule, so the coder is constant as well as the criterion. The rule: count an article when publicly accessible code appears, without executing it, to cover the article’s quantitative results, and the data are either included or obtainable by outside researchers through a documented procedure administered independently of the authors. Restricted data do not disqualify an article; discretionary data do. A German scientific use file with a published application route counts. Swedish register data reachable through Statistics Sweden counts. Partial coverage of a multi-study article does not count. Materials not publicly accessible do not count.

All eight 2023 positives survive. Eight of the eighteen 2025 articles qualify.

That is 32.0 percent in 2023 and 44.4 percent in 2025, against a reported rise from 32.0 to 66.7. Counting the one article whose repository required an access request produces 50.0 percent as an upper bound. Roughly two thirds of the reported increase follows from the change in what was counted.

The full case ledger, with the rule stated first, a resolving link and a check date on every entry, and the grounds for each downgrade named, is available separately. It was sent to Cohen before publication. He declined to review it.

The 2020 audit created a natural comparison, and Cohen created it without meaning to. He inspected four of the six issues of Volume 85. The two he skipped contain quantitative research articles that were never listed, never named, and never assigned a sharing status in public. Assignment to the treated and untreated groups depends on which issues he happened to sample, which has no obvious relationship to any author’s propensity to deposit. Same journal, same volume, same editorial regime, same year, same author pool.

Volume 85 contains twenty-one quantitative research articles. Fifteen were audited. Eleven of those were baseline non-sharers named in public. Five comparable articles sat in the unaudited issues. One of those five had already deposited and is not at risk. One further case, in the December issue, appeared close enough to the audit date that its exposure is ambiguous, and results are reported both including and excluding it.

The prediction, if public naming works through personal compulsion, is that named authors should subsequently deposit at a higher rate than unnamed ones.

The result is zero in both arms. No article in either group has a located first public deposit dated after November 30, 2020.

The second test uses the 2023 audit, where Cohen distinguished severity. Some articles were listed as supplying something partial. Others were listed as supplying nothing. Both groups were named in the same post. If naming operates through embarrassment, the second group had more to be embarrassed about. The result is again zero in both arms across the seventeen cases at risk.

The finding is that the outcome does not occur. Post-publication first deposit by ASR authors, across two cohorts and thirty-three articles, is close to a nonexistent behavior. Whatever a public audit accomplishes, it does not accomplish it by prompting authors to go back and post what they did not post before.

The design has no statistical power worth reporting and none is claimed. Eleven against four or five cannot detect an effect of any plausible size. What the comparison establishes is descriptive: in a hand-checked census of two full volumes, the behavior the intervention would have to change occurred zero times.

While cataloguing the corrections Cohen made to his audits, a pattern appeared that bears on all of this. The 2023 audit has been corrected four times in ways that changed an article’s classification. In each case an author of the article wrote in. In each case the material existed when he looked and his method had not found it. Simone Zhang and David Johnson’s Harvard Dataverse deposit was published in January 2023, a year before the audit. Christof Brandtner’s materials had been available throughout. Ferry and colleagues’ OSF project predated the check.

A fourth case was found not by an author but by this recount. Cheng and colleagues were classified as supplying nothing. A Harvard Dataverse record tied to the article’s appendix had been public since August 11, 2022 and the article did not link to it.

Four detection failures in a volume of twenty-five articles is sixteen percent. Three surfaced because a particular author happened to complain. The fourth surfaced because someone re-ran the audit. There is no reason to think the four are all of them.

The implication runs at this note too. Audits that follow an article’s links measure what an article’s links reveal. My own searches used titles, DOIs and author names against six repositories, and Cheng shows the method sometimes recovers unlinked material. It gives no basis for estimating how often it misses. The eight of eighteen reported above is a floor and the true 2025 figure is probably higher, as with every figure in every audit of this kind. This note rests on a comparison between two rules applied by one coder to the same articles, so a shared false-negative rate does not disturb it. Anyone reading the levels rather than the difference should read them as lower bounds.

The correction record is small. Across the three audits there are six documented corrections. Four changed a classification, and all four were prompted by an author of the paper being reclassified. One was self-initiated and arithmetic, catching an article counted twice. One came from a third party and fixed a broken link. No classification change originated with anyone other than an affected author.

The obvious reading is that authors are the people who notice, which is true and sufficient to explain the pattern. Six cases cannot distinguish that from anything else. What the record does establish is that the audit’s error-correction ran through a channel available to people with standing to write to Cohen and an interest in the outcome. Cheng did not write in and the misclassification stood for two and a half years.

Cohen’s 2020 post argues that sociology operates on trust because almost nobody checks anybody’s work. His own audits were checked by the authors they audited and by no one else.

The remaining route in Cohen’s framework is formal compulsion, which would provide an institutional explanation for any rise between 2023 and 2025. It does not apply. The American Sociological Review did not adopt a mandatory deposit policy in the period. It does not require a data availability statement. It does not verify submitted materials. The American Sociological Association’s ethical standards ask members to make data available to qualified researchers after publication, which is a professional norm rather than a submission requirement. Sage operates a tiered research-data policy and leaves the tier to each journal; Acta Sociologica and Sociological Methods and Research adopted stronger tiers, and ASR did not.

The internal evidence agrees. Three of the eighteen quantitative articles in Volume 90 carry no data availability statement of any kind. That is what an unenforced regime looks like from the inside.

So across the window studied here, the journal imposed nothing, verified nothing, and the twelve-point rise that survives a constant rule occurred without a mandate. Two of the three routes Cohen named are ruled out or produced no measurable events. What remains is normative drift, or ordinary variation. Twelve points across denominators of eighteen and twenty-five is inside the range that sampling alone would produce.

Three questions that seem open are open in a narrower way than they look. Whether deposit requirements raise deposit rates: they do. Olivia Miske and colleagues, coding the historical policies of 62 social and behavioural science journals across 2009 to 2018, found data available for 82 of 511 papers where no policy applied, or 16.0 percent, against 43 of 68 where data and code were required, or 63.2 percent. Their two remaining tiers, data alone and the full requirement including a pre-publication check, rest on sixteen papers and five papers, so the figures of 87.5 and 100 percent establish nothing about what those regimes add.

Whether verification raises reproduction: the evidence is too thin to say. In the same study, weighted precise reproduction runs at 40.7 percent where no policy applied, across 87 papers, and 75.7 percent where data and code were required, across 38. The two tiers above rest on fourteen papers and four, and give 70.5 and 65.0 percent. Nothing about what a pre-publication check adds can be recovered from four papers. These are observational associations across journals rather than estimates of what any policy caused, and the study’s headline rates, 53.6 percent precise and 73.5 percent at least approximate, are conditional on obtaining usable data, which the investigators managed for 143 of the 600 sampled papers.

The same study gives the first direct measurement of reproduction in sociology journals. Across the six in the frame, ASR, the American Journal of Sociology, Demography, Social Forces, the European Sociological Review and the Journal of Marriage and Family, fifty-nine papers were sampled and sixteen attempted. Two reproduced precisely. Four more reproduced approximately. Ten did not reproduce. ASR contributed ten sampled papers, seven never attempted; of the three attempted, one reproduced precisely, one approximately and one not at all. Nothing rests on three papers, and the counts are worth reporting because they are what exists.

Whether anticipation of inspection produces better submissions: it does not. The Quarterly Journal of Political Science has reviewed submitted packages in house since 2005, on the stated theory that transparency would motivate authors to be cautious, and its review checked that the code ran, that results could be located, and that they matched the paper. Reviewing twenty-four papers, Nicholas Eubank reported that twenty needed modification and fourteen produced results differing from the manuscript when the authors’ own code was run. He drew the conclusion himself: that these problems occurred despite authors knowing their code would be checked demonstrates the necessity of checking it. These are pre-publication states, caught and corrected, so the published articles are not defective; what the record shows is that anticipation was not sufficient. The exercise cost about six hours and roughly $180 per paper.

The American Journal of Political Science, whose verification policy is the most studied in the social sciences, processed 149 manuscripts and completed verification for 127 by February 1, 2018. Eight reproduced every analysis on the first try. Sociology has one comparable case: for the Fragile Families Challenge special issue at Socius in 2019, two of fourteen manuscripts reproduced on the first attempt.

The counterweight comes from Management Science, where Fišar and colleagues assessed 419 packages published under a mandatory disclosure policy and classified 95.3 percent as fully or largely reproduced, conditional on reproducibility being verifiable at all; among the failures, 88 percent turned on restricted data access. That is an assessment of packages that had already cleared an editorial completeness check, run after publication by outside reviewers, and it is not a first-pass rate. Its median reviewer spent four hours per package, against six for the QJPS execution check and eight for the Odum workflow that bundled archival curation with verification. The three numbers measure different work and should not be read as one service at three prices.

Sociological Science shows where sociology sits. It requires a replication package as a condition of publication and states in its reproducibility policy that it will confirm a paper lists its package contents correctly and will not perform replication tests. No sociology journal executes what it collects. Cohen’s December 2025 post discusses Sociological Science and package requirements without mentioning Eubank or the QJPS record, and I found no citation to Eubank anywhere in Cohen’s published work. Sociologists have engaged the paper elsewhere: James Moody and coauthors cite it in the Annual Review of Sociology, as do Jeremy Freese and David Peterson, and Bianca Manago cites it in The American Sociologist.

No study has publicly named individual authors for deficient sharing and then measured whether those authors subsequently deposited. The adjacent literature measures a different thing. Daniel Krähmer, Laura Schächtele and Andreas Schneck emailed 1,028 authors of papers using European Social Survey data and asked for their analysis code; 385 supplied it, or 37.5 percent. Claudia Acciai, Jesper Schneider and Mathias Nielsen, testing whether authors respond differently to different kinds of requester, approached 1,634 authors whose papers said data were available on request, using fictitious prospective doctoral students; 226, or 14 percent, either shared or said they were willing to. Private request, private reply, nothing public at either end.

The closest existing design is a randomized trial offering BMJ Open authors an Open Data Badge and measuring verified public deposit. Two of 54 in control, two of 57 in treatment. No effect. That tests positive recognition rather than public criticism, which are different treatments, so a null on one does not settle the other. It is the only randomized evidence on author-level reputational incentives and public deposit, and it points the same way as the zeros reported here.

Public naming works on institutions. When a news organization singled out universities and hospitals for failing to report clinical trial results, reporting rates at the named institutions rose from an average of 35 percent to 76 percent, and Stanford went from about a third to 85 percent after hiring six full-time compliance staff. Institutions have compliance offices, legal exposure and reputational management. Individual scholars have none of these.

The two comparisons here are underpowered by construction and are offered as bounded case series rather than hypothesis tests. The 2023 severity comparison carries a selection problem: articles in the partial group had already deposited something, articles in the nothing group had not, so the groups differ in the disposition the outcome measures and the confound runs against the prediction. Negative searches are the weakest form of evidence and the Cheng case demonstrates their failure mode. Where a deposit’s first public date could not be established from repository metadata or an archived capture, the case is recorded as unknown rather than resolved.

Repository states change. The pages that carry the most weight in the recount were submitted to the Internet Archive on August 26, 2026 and are cited by timestamped capture. One of them, an OSF project the audit had recorded as requiring an access request in December 2025, still redirected an anonymous visitor to a sign-in page eight months later. A GitHub repository named in the same audit stood at the same three commits. Nothing here can establish what any other repository looked like on any earlier date. The corrections catalogue rests on what Cohen published: post text, appended notes, and comment threads. Corrections made by email and never noted would not appear.

A public audit of a journal’s replication practices, repeated three times over five years by a prominent scholar, produced no located instance of an author going back to deposit. The journal it audited imposed no requirement and ran no check across the period. The rise the third audit reported is roughly two thirds an artifact of a widened category and, on a constant rule, is not distinguishable from ordinary variation. The audit’s own errors were corrected only by the authors it named.

The audits may have shifted expectations in ways no design here could detect, and the discipline’s conversation about transparency is louder than it was. The route Cohen named first, the individual scholar moved to act by seeing his paper listed, is not where the action is. Six hours and a couple of hundred dollars per paper buys a check that a decade of counting has not.

The case ledger for the recount gives the coding rule, twenty-six entries, resolving links, check dates and archived captures.

Posted in Sociology | Comments Off on What a Public Audit Does (V4)

Post: ‘Nathan Cofnas the Schlimazel’

Gimpel Neumann writes Aug. 22, 2026:

It has been one or two years since I first became aware of Nathan Cofnas. His first Substack piece did not appear until January 13, 2023—a paid discussion with Alex Kaschuta, followed by a few pieces that April. It was not really until January 2, 2024, however, that his writing took off, with the deliberately trollish “Why We Need to Talk About the Right’s Stupidity Problem.”

Nathan Cofnas was born in Chicago, sometime in 1986 or 1987 by most estimates. His first CV-listed publication, “Science Is Not Always ‘Self-Correcting’: Fact-Value Conflation and the Study of Intelligence,” appeared in Foundations of Science 21, no. 3, in 2016. Plenty followed: “Judaism as a Group Evolutionary Strategy: A Critical Analysis of Kevin MacDonald’s Theory,” in Human Nature 29, no. 2; “Is Kevin MacDonald’s Theory of Judaism ‘Plausible’? A Response to Dutton,” in Evolutionary Psychological Science 5, no. 1, in 2019.

Or take “How Gene-Culture Coevolution Can—but Probably Did Not—Track Mind-Independent Moral Truth,” published in The Philosophical Quarterly 73, no. 2, in 2023.

These, however, I have not read. They are not how I know Nathan Cofnas.

I know Nathan as a Substack persona. I have read his writings from “Podcast Bros and Brain Rot” to “A Guide for the Hereditarian Revolution”. I have watched his interviews. I have followed, like everyone else, his battles against ivory academia. I even watched his unfortunate debate against Dave Greene. But something nags at me: why do I know this Cofnas rather than the other one?

Why do I not know the Cofnas of Biology & Philosophy or The American Sociologist? Why do I not know him as an academic Jew, by the strength of his publications? Why not as a worthy man—a man of journals, monographs, and textbooks? Why does Cofnas’s public fame seem to exist almost entirely outside the scholarship that was supposed to establish him?

Gimpel Neumann calls Nathan Cofnas (b. c. 1986) a schlimazel, a man on whom the soup gets spilled. The frame is warm, the mazel exposition is charming, and the conclusion is wrong. Cofnas read the reputation economy correctly, chose the branch with higher variance, and drew from the tail. He has agency. Ghent and Cambridge have responsibility. God can stay on His throne.

Cofnas published in Foundations of Science in 2016, in Human Nature and Evolutionary Psychological Science on Kevin MacDonald, in The Philosophical Quarterly in 2023 on gene-culture coevolution and moral truth. That record bought him an early career research fellowship at Cambridge and a research associateship at Emmanuel College in 2022. The system worked. It gave him what it gives: a term appointment, a hundred readers who could referee and cite him, and a line on a CV that opens the next term appointment. Neumann asks why the journals failed to make Cofnas famous. They never make anyone famous. That is not their function.

What the guild cannot deliver is a public, and a public is what Cofnas wanted, because a public is what his theory of change requires. He said so in the January 2, 2024 essay. The right needs to challenge the equality thesis. Right-wing intellectuals should disseminate accurate information about race differences and build a political philosophy attractive to current left-wing elites. Whatever one thinks of that program, it cannot be executed in Biology & Philosophy. The people it aims to convert do not read Biology & Philosophy. So he wrote for the people who read Substack, and he wrote in the register that gets read there, and he titled the essay to be shared by people who were annoyed by it.

The trollish tone was a choice. Cofnas is a philosopher of science who spent years learning to write for referees. A man who can write for The Philosophical Quarterly and then writes “Why We Need to Talk About the Right’s Stupidity Problem” has made a decision about audience. The decision paid. Within two years he had a readership, an income stream, invitations, and a name that people in three countries recognize. He also had Emmanuel College ending its association with him in 2024 after complaints about a post arguing that Black students would nearly vanish from Harvard under colorblind admissions. Cambridge’s own review found no discrimination or harassment and no breach of speech rules. The association ended anyway. That is the posted price, and he had watched enough predecessors pay it to know the figure.

Ghent settles the question. When his hire became public in March 2026, nearly a thousand people signed a petition demanding it be stopped, and more than three hundred staff and students, including forty-eight members of his own department, three deans, and two former deans, signed a letter calling his views incompatible with the university’s ethics code. He took the job. Both parties knew what they were getting. Then he went on a podcast and said he was continuing his work on race realism and the hereditarian revolution while employed at Ghent, and Ghent’s legal office quoted that statement back to him in the letter of August 20, 2026, opening a preliminary disciplinary investigation and suspending him. His research mandate, they wrote, covers the future of liberalism and individual freedom and equality, and race realism falls outside it.

A man who takes a post over the objection of a third of his department, then announces on a podcast that he is continuing the work they objected to, is not a man to whom things happen. He is a man running a strategy. The strategy has a name in finance. He is long volatility. He accepts a high probability of small losses and periodic institutional expulsion in exchange for a small probability of a very large payoff, which in his case would be the collapse of a taboo he believes is empirically false and socially catastrophic. Bets like that lose most of the time. Losing most of the time is priced in.

July and August delivered the payoff and the tail in the same motion. Cofnas ran Jason Arday’s 2015 doctoral thesis through plagiarism software and published the matches: dozens of passages lifted verbatim or near verbatim from Paula Zwozdiak-Myers’s 2009 Brunel thesis, much of it uncited at the point of copying, along with problems in four of ten papers drawn from Arday’s own list of principal publications. Times Higher Education had held the same story since September 2025, having spiked it after Arday engaged Carter-Ruck. Once the screenshots were public the suppression failed. The Telegraph commissioned an analysis that found more than a hundred identical or near-identical passages, a later count reached 188 sentences, and academics with no interest in Cofnas’s racial politics examined the two texts and reached the same conclusion. Cambridge, which had called the critics racist, opened an investigation. Arday resigned on August 5 and was found dead at his home in London nine days later, at forty-one. Then the suspension.

Any strategy built on public accusation carries one tail risk above all others, which is that the target is a person, and persons break. That risk was known before the first post went up. Under Neumann’s frame, a man’s death enters the ledger as soup on Cofnas’s trousers. Under the agency frame, it enters as the realized cost of a chosen method.

The luck frame also fails on its own terms. A schlimazel gets no cavalry. Cofnas got one in four days. By August 24 more than seven hundred academics had signed an open letter to rector Petra De Sutter (b. 1963) calling the proceedings retaliation for publicizing misconduct, among them David Deutsch (b. 1953) and Nigel Biggar (b. 1955) and fifteen Oxford signatories and two dozen business school professors. That is the asset the strategy was built to accumulate, arriving on schedule. He spent two years constructing a constituency outside the university so it would exist on the day the university moved against him. It exists. Whether it is large enough to hold Ghent is the question, and it is a question about coalitions.

Ghent hired a man for his work on liberalism knowing he was famous for something else, and now proposes to discipline him because he is famous for something else. The mandate argument is the useful part of the letter. It lets a university avoid saying that it punishes views, and say instead that the views fall outside the contract. Any academic whose public writing exceeds his grant description is exposed to that argument. Ghent has built a tool with wide application, and other rectors are watching to see whether it holds.

So the pattern in Cofnas’s career is a man who concluded that scholarly credentials no longer convert into influence, that the public branch of the reputation economy is where beliefs now move, and that the price of entry is his standing in the guild. He was right about the split. He priced the ticket correctly. The bet is still running.

Posted in Nathan Cofnas | Comments Off on Post: ‘Nathan Cofnas the Schlimazel’

Philip N. Cohen and the Sociology of the Denominator

Philip N. Cohen’s career can look unusually scattered if you read it by subject. He has written about women’s suffrage, racial inequality, housework, occupational segregation, cohabitation, marriage, divorce, same-sex parenting, genetics and race, fertility, COVID-19, peer review, preprints, open science, public sociology, and the politics of falling birth rates. His own bibliography now divides naturally into two large bodies of work, one on families and inequality and another on scholarly communication and science. But after reading the corpus chronologically, I think that division misses the strongest continuity.

Cohen has spent thirty years doing a sociology of the denominator.

By denominator I mean the rule that decides who has been counted and what has been grouped with what. Cohen’s characteristic move is to encounter an apparently stable social fact and ask what the rule was. Change the universe, the reference period, or the boundary of the category, and the fact may change. Two related habits grow out of that one, and I take them up later: a suspicion of explanations pitched at the level of the individual, and a suspicion of trends converted into directions of history.

Cohen’s first journal article, the 1996 “Nationalism and Suffrage: Gender Struggle in Nation-Building America,” suspects a category that looks more universal than it is. “Women’s suffrage” sounds like the political inclusion of women. Cohen instead reconstructs a movement in which leading white suffragists increasingly made their claim through white citizenship, national strength, racial hierarchy, and an alliance with white men. The category “women” concealed conflicting interests among women themselves. White suffragists could advance their own political status while helping exclude Black women and men from the polity. Cohen ends by insisting that feminist scholars treat those women as agents making purposeful choices.

Three years later, with Suzanne Bianchi, the same intellectual operation becomes statistical. “Marriage, Children, and Women’s Employment: What Do We Know?” begins from an apparently straightforward question: how much do married mothers work? But the Current Population Survey permits different reference periods and different universes. Cohen and Bianchi show that estimates of women’s full-time employment change substantially depending on which is used. Welfare policy was being justified partly by assumptions about how much mothers normally work. If the denominator changes the normal rate, it changes the policy comparison.

The following year, Cohen and Lynne Casper made the same kind of intervention in “How Does POSSLQ Measure Up?” The issue was cohabitation. A government measure that appeared to count unmarried partners was an operational construction whose performance could be tested against alternative ways of identifying couples.

That pattern runs through Cohen’s early inequality research. In “Individuals, Jobs, and Labor Markets: The Devaluation of Women’s Work,” with Matt Huffman, the wage attached to female-dominated work cannot be understood simply by comparing individuals. Jobs are situated within labor markets, and those contexts alter the relationship. In “Racial Wage Inequality: Job Segregation and Devaluation Across U.S. Labor Markets,” segregation and devaluation are analytically separated. In “The Gender Division of Labor: ‘Keeping House’ and Occupational Segregation in the United States,” the usual segregation statistic changes once women doing unpaid housework are restored to the division of labor.

The cumulative lesson is that inequalities attributed to persons are often properties of arrangements around persons. Households, jobs, establishments, occupations, labor markets, and policy regimes keep appearing between the individual and the outcome. Cohen’s 2007 work with Huffman on “Black Underrepresentation in Management across U.S. Labor Markets” asks why establishments operating in areas with larger Black populations are more likely to underrepresent Black workers in management.

Family research remains the largest substantive part of Cohen’s career, but much of it is an argument against treating “the family” as the natural unit at which family outcomes should be explained. Extended households change the employment possibilities of single mothers. National policy changes the relationship between individual bargaining resources and housework. Marriage markets differ by race. Economic shocks affect divorce and fertility. The family is nested inside structures that alter what family means.

Cohen’s second recurring suspicion is directed at trends that are turned too quickly into historical directions.

The titles tell part of the story. “Stalled Progress? Gender Segregation and Wage Inequality Among Managers, 1980–2000.” “Headed Toward Equality? Housework Change in Comparative Perspective.” “The ‘End of Men’ Is Not True.” “The Coming Divorce Decline.” “What’s the Story? Family Demography at the End of Progress.” “Rethinking Marriage Metabolism.”

Cohen encounters a real change and objects to the story attached to it. Women are advancing, therefore men are ending. Housework is becoming more equal, therefore equality is the destination. Marriage is declining, therefore marriage is disappearing. Fertility is falling, therefore population catastrophe is approaching. Cohen distrusts the conversion of movement into destiny.

When Hanna Rosin claimed that “young women” were outearning young men, Cohen unpacked “young women” and showed that the statistic described single, childless women under thirty who lived in metropolitan areas and worked full-time and year-round. Rosin said women had become a majority of managers. Cohen noticed that the underlying category was “managerial and professional,” which combined managers with female-dominated professions such as nursing and elementary education. The category created the headline.

Cohen begins that essay with a sentence that could serve as a manifesto for his career: statistics are not “mere technical details.” They are attempts to represent reality. That is why the boundaries matter.

The anti-teleological position becomes explicit in “What’s the Story? Family Demography at the End of Progress.” Cohen argues against imagining normal modernity as a track along which societies travel while pandemics, climate change, inequality, identity conflict, and political disorder interrupt the journey. Those things may be constitutive features of the world whose demography we are trying to explain. Even the second demographic transition begins to look suspicious if it is treated as another version of history moving in an incontrovertible direction.

This helps explain Cohen’s more recent interest in falling fertility. His 2025 article, “Negative Views of Falling Birth Rates in the United States Come Mostly from the Right Wing,” and his continuing work on population panic fit the older pattern. The subject has changed. The argumentative opponent has not. A demographic trend is being recruited into a large story about social decline. Cohen’s instinct is to separate the measured change from the narrative laid over it.

Return now to the first habit, because it is the one that carries him into a new field.

“Women.” “Managers.” “The Black middle class.” “Cohabitors.” “Married people.” “Young women.” “Science.” “Peer review.” Each can operate as a useful category. Each can also conceal enough heterogeneity to make a claim true only because unlike things have been pooled.

That is what makes the later open-science work a logical extension of his interests.

The 2022 collaborative paper “PReF: Describing Key Preprint Review Features” begins from a classification problem. “Preprint review” sounds like a recognizable practice, but the activities collected under that term differ in who initiates review, how reviewers are selected, whether identities are known, whether competing interests are checked, whether authors can respond, and what outcome the review produces. The solution is to decompose the label into standardized features so different systems can be compared on common dimensions.

That is Cohen’s old move in a new field.

In 1999, “women’s employment” had to be decomposed by universe and reference period before a trend could be interpreted. In 2013, “young women” had to be decomposed before Rosin’s claim could be evaluated. In PReF, “peer review” has to be decomposed before anyone can compare review systems or measure their evolution. The substantive distance between those projects is enormous. The methodological distance is small.

The same continuity appears at the level of causal explanation. Around 2018, Cohen’s bibliography begins to develop a second program alongside family demography: open scholarship, SocArXiv, peer review, scholarly infrastructure, and the production of knowledge. He never stops doing demography. His recent work includes marriage expectations, marriage events, fertility, mortality, and birth rates at the same time that he is publishing about scholarly communication.

The new program again moves the explanation away from individual behavior and toward context. In “The Scholarly Knowledge Ecosystem: Challenges and Opportunities for the Field of Information,” Micah Altman and Cohen define the scholarly ecosystem broadly enough to include stakeholders, markets, policies, organizations, norms, infrastructure, and educational systems. Scientific reliability is an outcome produced inside an institutional ecology.

The analogy with the earlier labor-market work fits. An employee’s outcome cannot be understood without the job, establishment, occupation, and labor market. A scientific paper cannot be understood without journals, repositories, incentives, review systems, publishers, universities, funders, and technical infrastructure. In both cases Cohen resists explanations that stop at the individual because the relevant machinery is operating one or more levels above.

There is a political continuity too.

Cohen’s 1996 suffrage paper is a critical account of white feminism, nationalism, and racial exclusion. His 2007 review essay “Confronting Economic Gender Inequality” is hardly value-neutral. What changes is that Cohen becomes increasingly explicit about what political commitments mean for epistemology.

The pivotal document is probably his 2015 “How Troubling Is Our Inheritance? A Review of Genetics and Race in the Social Sciences.” Cohen argues against Nicholas Wade’s attempt to connect population genetics to racial differences in social behavior. He says that he sets the evidentiary bar high for racial-genetic explanations because the consequences of a false positive are potentially severe. He calls this a political judgment and compares it to the asymmetric treatment of error in criminal justice and medical research.

That passage makes it difficult to describe Cohen’s later open-science work as an attempt to construct a value-free machine for deciding truth. He does not believe such a machine exists. Altman and Cohen’s scholarly-ecosystem work argues for information systems aligned with human values. The point of transparency is to make judgment inspectable.

That position is developed in Cohen’s 2025 book, Citizen Scholar: Public Engagement for Social Scientists. Its chapter sequence moves from description through open scholarship and peer review to social media, activism, and citizenship. The book’s argument is that scholarship and citizenship can be integrated, provided the scholar remains answerable to evidence.

Seen chronologically, there is therefore a critic-to-builder trajectory in Cohen’s career. By the early 2010s he is increasingly visible as an auditor of public and scholarly claims. Rosin’s statistics. Mark Regnerus’s same-sex parenting research. Nicholas Wade’s racial genetics. Alice Goffman’s methods. Claims about poverty, marriage, and family decline. The scholar stands outside the claim and asks whether it survives inspection.

Then Cohen becomes involved in building the infrastructure through which claims circulate. SocArXiv launches in 2016. He works on scholarly communication, preprints, review taxonomies, openness, editorial boards, and research infrastructure. The person who spent years inspecting products coming out of the scientific machine increasingly starts working on the machine.

His own scholarly-communication research states the methodological principle. In “Interventions in Scholarly Communication: Design Lessons from Public Health,” Altman, Cohen, and Jessica Polka argue that scholarly reform suffers from an absence of standard measurement, systematic data collection, and comparable evaluation. They mock the common intervention strategy as “Do it now, check sometime.” Their proposed remedy is to identify common measures and outcomes, build infrastructure for tracking them, incorporate assessment into interventions from the beginning, and create consistent longitudinal evidence.

PReF makes the same point in another form. If heterogeneous review practices are going to be compared across time, first define stable features that distinguish them. Only then measure the trend.

Public argument rewards compression. “ASR went from one-third to two-thirds” is intelligible in a way that a five-category taxonomy of data accessibility is not. A scholar trying to participate in public life must turn a complicated empirical world into categories, numbers, sentences, and stories. Yet Cohen’s scientific work repeatedly demonstrates that the compression is where trouble enters.

This tension sits near the center of Citizen Scholar. Cohen wants scholars to leave the seminar room and enter public argument without surrendering the standards that make scholarship worth hearing. But his own career suggests that the two activities place different pressures on description. The public intellectual needs a story. The empirical sociologist keeps discovering reasons the story is too clean.

His three informal audits of replication materials at the American Sociological Review are where that pressure becomes visible and countable, and I have taken them up separately.

That may be the most interesting trajectory in the entire corpus — a progression from studying how social arrangements produce observable outcomes to studying how knowledge arrangements produce observable facts, followed finally by an attempt to live inside both systems at once as a citizen scholar.

The objects keep getting larger. Women within families. Workers within jobs. Jobs within establishments. Establishments within labor markets. Families within policy regimes. Papers within journals. Scholars within an information ecosystem. But Cohen keeps asking variants of the same question: what structure surrounding the observation made it look the way it does?

And alongside that question is another: what disappeared when we named it?

“Women” can hide racial conflict. “Young women” can hide a highly selected demographic subgroup. “Managers” can become managers plus professionals. “Peer review” can hide radically different institutional processes.

Cohen’s sociology is strongest when he refuses to let the noun settle the empirical question.

That is why denominator skepticism seems a better description of his method. Before accepting a social fact, ask who counts. Ask what the comparison class is. Ask whether the category means the same thing in every observation. Ask whether a measured difference has been converted into a causal story, or a trend into a theory of history.

Those habits connect work that otherwise looks unrelated.

Cohen argues that statistics matter enough that their construction has to be exposed. Categories can mislead because reality exists outside them. A better denominator can produce a better description.

That faith in better description survives into the metascience work. The solution to unreliable scholarly communication is to expose provenance, make evidence inspectable, improve measurement, clarify categories, build institutions that permit criticism, and give outsiders enough information to decide whom to trust.

There is a revealing passage in Cohen’s 2023 chapter “How Do We Tell What’s True?” in which he refuses to put conventional peer review on a pedestal. Expert credentials, institutional reputation, disclosed political commitments, linked evidence, open criticism, and transparent sourcing all provide imperfect signals. Knowledge becomes more trustworthy when claims can be traced and interrogated.

That is also a description of Cohen at his best.

His enduring subject is the distance between a claim and the machinery that produced it.

For thirty years Cohen has been opening that machinery.

Andrew Gelman is the nearest single analogue. Substantive researcher, then methods critic, then a daily blog auditing other people’s published claims, then a participant in reforming what the field accepts.

The three acts match. Substantive researcher, then critic of method, then a public auditing operation, then a hand in what the field will accept. Cohen’s blog has run since 2008 and Gelman’s since 2004, and both use the same instrument: take a published claim, go to the numbers behind it, show the reader the arithmetic, name the paper and the author. Both write for a mixed audience of colleagues and journalists. Both get called a bully by people they have corrected.

Two differences. Gelman reruns. When he takes apart a paper he pulls the data where he can and does the analysis, and his commenters do it too, and the corrections that come out of that blog are corrections about what the numbers show. Cohen inspects. His audits ask whether a package exists and what it contains, and he says he does not run the code. That is the difference between an auditor and an inventory clerk, and it explains why his findings are about availability while Gelman’s are about results.

Gelman also audits his own side. His targets are chosen by whether a claim looks statistically implausible, and the record includes his own papers, his friends’ papers, and work whose conclusions he likes. That indifference is his central asset.

Kieran Healy holds Cohen’s epistemology. His work argues that the measure and the rendering are constitutive of the finding, and his 2017 essay against nuance argues that sociology overvalues theoretical elaboration and undervalues clear description. That last point is Cohen’s descriptive turn stated as a general position about the discipline, published before Cohen said it at Wisconsin, by a sociologist Cohen surely reads.

Healy built no audit. No annual counts, no named lists, no preprint server, no correction demands. He wrote a book about how to draw graphs and taught a generation of sociologists to think about rendering.

So the epistemology does not produce the enforcement. Two men hold roughly the same view of what a measure is and only one of them audits. Whatever explains Cohen’s project, it is not the belief that categories manufacture findings, because Healy holds that belief and does something else with it.

Jeremy Freese occupies the position in the discipline that Cohen might have taken. He is sociology’s designated metascience thinker, coauthor of the field’s review article on replication, a writer on preregistration and research transparency. He works through the literature rather than through counts and lists.

Freese and Peterson cite Nicholas Eubank. Cohen does not. The sociologist doing metascience as scholarship engaged the one journal that ran the experiment; the sociologist doing metascience as enforcement did not.

Otis Dudley Duncan (1921-2004) built the tools. The dissimilarity index he developed with Beverly Duncan became the standard measure of residential segregation, and he brought path analysis into sociology. Then he spent his later career arguing that sociology’s measurement was pre-scientific, that its quantities frequently did not measure what their names claimed, and that the discipline had adopted numerical machinery without asking what the numbers meant. His late book on social measurement is a sustained warning against exactly the pooling Cohen keeps undoing.

Duncan questioned the segregation index he had built and the causal modeling he had introduced. Cohen has done that with PReF, where he decomposed a category to build a measurement tool. He did not do it with the audit categories. Duncan is the standard, and the standard is set inside Cohen’s own lineage: this is a demographer working on segregation, which is Cohen’s own subject, who made auditing his own instruments the last act of his career.

Barry Wimpfheimer’s argument is that a category which arrived later, legal-analytic reading, has governed how earlier material gets read, and that the material looks different once you stop letting the later category set the terms. That is the same move on texts rather than on variables, which tells you the operation is a habit of asking what the naming did.

What the set adds up to is a way of locating Cohen.

The category operation is not his: Healy has it, Duncan had it, Wimpfheimer has it in another material. The three-act trajectory is not his: Gelman ran it a decade earlier in an adjacent field. The metascience is not his alone: Freese has been writing it in sociology throughout.

What is distinctive is the combination plus the enforcement. Cohen alone among these men publishes lists of named papers with a verdict beside each, repeatedly, over years. That is the residue after the comparisons subtract everything shared.

Posted in Sociology | Comments Off on Philip N. Cohen and the Sociology of the Denominator

Measuring the Same Thing Twice (v3)

Philip N. Cohen has inspected the American Sociological Review three times to see whether its authors post the code and data behind their published results. The first inspection came in November 2020, the second in January 2024, the third in December 2025. Each time he read through the quantitative articles, looked for links to materials, and counted. He did not run any of the code. He says so.

The third post opens by reporting that the journal is making progress. Two-thirds of the quantitative articles in Volume 90 provided replication materials, he writes, against roughly a third in the previous audit and roughly a quarter in the first. Twelve of eighteen, up from eight of twenty-five.

The trouble is that the two numbers count different things.

In January 2024 Cohen sorted twenty-five quantitative articles into three columns. Eight had what he called a full package, by which he meant “data and code sufficient to replicate the papers.” Five or six others, depending on whether you read his prose or his table, supplied something less: partial data availability, a set of variable codes, fragments. The rest supplied nothing. The distinction between the first column and the second is the entire architecture of that audit.

In December 2025 the columns merge. His positive category now holds articles that “provide packages of code and data, or at least information about getting the data.” He adds that he counts materials parked on personal websites, which is not where replication files belong, and excludes materials described as available on request. Twelve of eighteen articles clear that bar.

He states the new rule in the post, in the sentence quoted above, for anyone who reads it. What he does next is compare the resulting figure against numbers produced by the older, narrower rule, and read the difference as movement in the journal.

I recoded the eighteen 2025 articles under the 2023 standard, one fixed rule, no code executed, matching his method. The rule: count an article when publicly accessible code appears to cover its quantitative results and the data are either included or obtainable by outside researchers through a documented procedure that does not run through the authors. Restricted data do not disqualify an article. A German scientific use file with a published application route counts. Swedish register data reachable through Statistics Sweden counts. Data that arrive only if the authors decide to send them do not.

Under that rule, eight of the eighteen 2025 articles qualify. Forty-four percent, against thirty-two percent in 2023.

So the journal improved, but the improvement he reports runs from thirty-two to sixty-seven percent, a gap of thirty-five points, and the improvement a constant rule can find runs from thirty-two to forty-four, a gap of twelve. Roughly two-thirds of the reported rise comes from the widening of the category.

Four articles produce the difference, and none of them is bad work. Fabiana Silva, Irene Bloemraad and Kim Voss deposited materials for “Frame Backfire” at the Open Science Framework, and Cohen’s own table records that the project required an access request when he checked it on December 9, timestamped to the minute. He counted it anyway, and elsewhere in the same post tells the journal that permission-gated projects should not be allowed. The barrier is still up. A reader who follows the article’s link today without an OSF account arrives at a sign-in page. Counting it regardless gives nine of eighteen, the outer bound of the recount. Mabel Abraham, Tristan Botelho and James Carter posted data and code covering the main results of two of the studies in “(Not) Getting What You Deserve”; the article contains more than two. Daniel Scott Smith and coauthors, writing on “How Values and Uncertainty Shape Scientific Advance in Peer Review,” posted a repository whose own documentation says it holds no data, excludes some code, and cannot reproduce the machine learning benchmarks without material the authors supply on request. Eight months after Cohen named it, the repository stands at the same three commits. David Brady, Aliza Luft and Ezra Zuckerman Sivan, reanalyzing another team’s work in “How Does Culture Matter for Attainment, and How Would We Know If It Did?”, posted a PDF of their code, which states that the underlying data came from the scholars whose work they were reanalyzing and cannot be passed along. It also carries author-specific file paths and merges against a dataset it does not supply, so an outside reader could not run it even holding the data.

The obvious defense is one Cohen has already made. In 2024 he described his own audit as a tiny study of poor quality. In 2025 he calls his methods “squishy.” A man who labels his instrument that way has told his readers what weight to put on it. The answer is the first line of the same post, where he writes that the journal is making progress and sets the three figures against each other to show it. The disclaimer and the directional claim sit in the same paragraph.

Cohen has spent thirty years teaching other researchers not to do this.

The operation runs through his corpus from the beginning. In 2000 he and Lynne Casper built a new historical measure of cohabitation because the standard one misclassified people and distorted the trend it was supposed to reveal. In 2002 he showed that part of the apparent decline in the marriage premium for men came from cohabitors sitting inside the never-married comparison group, and that pulling them out shrank the decline. He and Lynne Casper split multigenerational households into hosts and guests because knowing that someone lives in an extended household tells you nothing about their position in it. He and his coauthors argued that a Black middle class defined through married households misses the growing population of never-married professionals living alone. He recoded keeping house as an occupation and found that women leaving unpaid domestic work accounted for as much of the decline in occupational segregation as desegregation within paid work. With Matt Huffman he took the established association between local Black population share and racial wage inequality and split it into two candidate processes, segregation and devaluation, and reported that one survived the data and the other did not.

A category that pools things behaving differently will manufacture a finding. Three cases show that he knows this. The first is PReF, the preprint review framework he built with collaborators, a fourteen-author project from 2022. It begins from the observation that peer review names procedures too different to compare, and replaces the label with eight standardized descriptors so a reader can see what a given review consisted of. The point of the exercise is to make comparison possible by first fixing what is being compared.

The second is his 2013 essay “The End of Men Is Not True: What Is Not and What Might Be on the Road Toward Gender Equality” on Hanna Rosin’s The End of Men, which is combative from the first page. The claim that young women out-earn young men turns out, on inspection, to describe single childless full-time metropolitan workers, and Cohen walks through the restriction before reanalyzing the narrowed group and finding that men still come out ahead within education categories. He does the same to her claim about managers, which rests on a category folding female-dominated professions into a statement about management. Polemic did not cost him the method. He goes claim by claim, category by category, for pages.

The third is closest to home. In 2023, with Micah Altman and Jessica Polka, he published “Interventions in Scholarly Communication: Design Lessons from Public Health” in First Monday. The complaint there is that open science reforms get launched without stable measures or systematic evaluation. The prescription is common measures, assessment designed in from the start, and comparable outcomes tracked over time. That appeared two years before the 2025 audit, addressed open science reform specifically, and carried his name second.

So the ASR audits are a case of a technique he had, used under fire, and prescribed in print for this exact class of work, and then did not apply to a measurement he was building, in the one place where the measurement was going to be read as evidence of a trend.

Applying his 2023 standard to his 2025 cases moves sixty-seven percent to forty-four. The journal improved. It improved by about a third as much as reported, and what survives a constant rule is small enough that eighteen articles cannot separate it from ordinary year-to-year variation.

Cohen has spent years asking sociologists to correct the record when the record is wrong, and has corrected his own audit twice when authors wrote to tell him he had missed their packages. I sent him the coding rule and the twenty-six cases, with links and dates, before publishing this. He declined to read it. The ledger stands open, and if any of it is wrong, it will be fixed here. (V3)

Posted in Sociology | Comments Off on Measuring the Same Thing Twice (v3)

Rebecca Sear: ‘National IQ’ datasets do not provide accurate, unbiased orcomparable measures of cognitive ability worldwide’

Here is her paper, which is strongest where it is most boring, and weakest where it is loud.

Using Lynn and Becker’s own sample-type coding against them is the move that will survive. By their own generous definition, where 74 orphans in a single Eritrean orphanage counts as “national,” only 34 percent of the 656 samples qualify. That number comes from their spreadsheet. A defender has to attack his own data to attack it. The same goes for the case catalogue: Angola from 19 malaria-free individuals, Somalia from refugee children in Kenyan camps, Botswana from 140 adolescents recruited in South Africa on ethnic-proxy grounds. These are checkable in an afternoon.

Extending the catalogue to Denmark, Norway, Sweden, and Ireland closes the standard escape route. If Denmark’s number comes from fifth graders measured in 1968 and Norway’s from a cod liver oil trial and an epilepsy study, then the problem is the method.

Two arguments do most of the intellectual work. The first is non-independence. Twenty-five percent of current values are averaged from up to three neighboring countries, then fed into regressions that assume independent observations. That invalidates the published literature on statistical grounds alone. It is the argument an econometrician has to answer. She gives it a paragraph.

The second is the stability. The sub-Saharan average has stayed between 67 and 70 across four versions built from substantially different samples. Sample composition churns; output holds. That pattern is what a target value looks like. She gives it one sentence inside a longer passage about Wicherts. I would build the paper around it.

The fraud claim is where she overreaches. She makes it twice, in the introduction by way of Wicherts and in her own voice near the end. What the evidence supports is a research program with an announced prior, documented ad hoc adjustments toward that prior, and no correction after twenty years of criticism. Motivated construction. Fraud requires intent to deceive, and she has not established it to a standard that survives a hostile reading. The cost is practical. Say “fraudulent” and the defenders litigate the word instead of the sample sizes, and she is at a British university writing about a man whose institute still has friends.

She also gives the opposition its weakest case. Warne and Kirkegaard’s argument is convergent validity: national IQ correlates high with PISA, TIMSS, and adult literacy measures. A normal person looks at that and recognizes truth. Her answer is that the correlation tells us nothing. That is a weak argument. Random measurement error attenuates correlations toward zero. If the inputs were as noisy as she describes, the series should correlate with nothing. It correlates with a lot.

If comparable cross-cultural measurement is impossible in principle, then Wicherts’s corrected figure of 80 for sub-Saharan Africa is also meaningless, and so is the claim that Lynn biased the numbers downward, which presupposes a truer value. She wants the practical argument and the in-principle argument at once. They pull in opposite directions. The practical argument stands alone and I would let it. Measurement invariance is testable, and failures of it are findings rather than axioms.

Wicherts is doing heavy lifting for her as an ally while holding a position she rejects. His 80 is still low. His conclusion was that national IQs fail to support evolutionary explanations, and he thinks a regional estimate can be made and made better.

Slobodian (b. 1978) on Lynn’s influence is a claim about political history rather than about data quality. Including it tells you the target is journal editors and integrity officers rather than fence-sitting quantitative social scientists. That may be the right call given her goal, which is retraction. It makes the piece easier to file under politics.

The buried story is in the footnotes. Comparative Sociology published Jensen and Kirkegaard after the editor was told about the dataset. Nature Sustainability published Chu et al. in 2026 after Springer’s research integrity group was told. Two named journals, alerted in advance, publishing anyway. That is a gatekeeping story with dates and actors.

Small things. The DSM is the American Psychiatric Association, not the American Psychological Association. Warne appears as 2023 in text and 2022 in references. “a prior” for “a priori.” “the wholly inadequate of the primary data sources” in the conclusion is missing a word. She excludes Lynn for keeping control groups and for keeping treatment groups; both criticisms are fair, but stated together they read as scoring both ways.

The structural parallel to Jason Arday. Documented methodological failure, institution notified, institution proceeds, critic escalates from method to misconduct. The escalation is what happens when the first route fails.

Posted in IQ | Comments Off on Rebecca Sear: ‘National IQ’ datasets do not provide accurate, unbiased orcomparable measures of cognitive ability worldwide’

The More I Love, the More Careful I Am

When my life overflows with people I love, I tend to be circumspect. When my life lacks love, I tend towards shock and awe.

Bacon called us lovers hostages to fortune. A man with a marriage, children, a congregation, an income, and friends who can be turned against him speaks with one eye on the people who can take those things.

Caution tracks exposure. What counts is how much of what you love sits in an unsafe place. A tenured professor has a great deal and risks little, because the thing he most values cannot be removed by an angry reader. An anonymous account can have a family, a mortgage, and a security clearance and still post anything, because nothing traces back. A wealthy man with no employer and no board is in the same position as a man with nothing, though for opposite reasons. Possession without traceability produces no restraint at all.

Then there is the case where recklessness is the asset. Some men are paid for it. The audience arrives for the transgression and leaves when it stops. For that man the precious thing and the reckless speech are the same object, and caution costs him everything.

The causal arrow also runs backward. Reckless speech destroys the precious things, and the man who has lost them speaks with less restraint the following year. Divorce, firing, expulsion from a community, and then a freer tongue. The correlation I observe in a given moment may be the wreckage of earlier speech. This produces a ratchet. Each round of loss lowers the price of the next round.

And a great many reckless posters have plenty to lose and believe themselves safe. Small audience, private account, friendly comment section, an assumption that no one screenshots. The recklessness comes from a bad estimate. Anyone who has watched a man lose a career over a post knows he did not think he was gambling.

A father who thinks his children will inherit a worse world, a member of a community who thinks the community is destroying itself, a man who watched his father expelled over a doctrinal fight. For him silence is the thing that costs. Attachment produces speech.

So the generalization survives: Restraint rises with the amount of loved and vulnerable property that lies within the punishing reach of your audience, discounted by the accuracy of your estimate of that reach.

Two hazards. Used on other people it becomes a cheap discount on unwelcome speech. He only says that because he has nothing, therefore ignore him. This would dismiss every prophet, and prophets are the test case, since the standard biblical pattern is a man who has something and forfeits it.

The precious things give you a permanent, respectable, sincere reason never to say the difficult thing. You cannot tell the two apart by introspection. The only test I know is whether the things you decline to say cluster around a single group of people who could hurt you. If the silences are distributed, it is judgment. If they all point one direction, it is fear.

Posted in Blogging | Comments Off on The More I Love, the More Careful I Am

A Passage-Level Count of King’s Dissertation, and What It Says About the Arday Case

Two men submitted doctoral theses sixty-one years apart. Both were plagiarized. Both universities were told. Neither university revoked the degree. Each man went on to great success.

Martin Luther King Jr. submitted “A Comparison of the Conceptions of God in the Thinking of Paul Tillich and Henry Nelson Wieman” to Boston University on April 15, 1955 (audit). Jason Arday (1985-2026) was awarded a doctorate in education by Liverpool John Moores University in 2016. In 2023 Cambridge appointed him professor of sociology of education, the youngest Black professor in the university’s history. He resigned on August 5, 2026, hours after Cambridge announced an investigation into his academic qualifications and honorary appointments. He was found dead at his London home on August 14.

Start with the documents.

I (meaning ChatGTP and Claude) ran a passage-level count on King’s thesis using the King Papers Project edition, which prints the parallel source text under every borrowing its editors identified. Four hundred fifty editorial notes. About 23,500 words of printed parallel against roughly 73,500 words of body text, which is a third of the thesis. Chapter III on Tillich runs the heaviest. Chapter VI, the conclusions, carries none. Two hundred fifty-three of the notes show a run of ten or more consecutive words shared between King’s page and the source. Jack Boozer’s 1952 Boston University dissertation, “The Place of Reason in Paul Tillich’s Concept of God,” appears in eighty notes; King names Boozer on twelve of those pages. David Roberts appears in thirty notes and King never names him on the page. John Herman Randall Jr. appears in twenty-one and is named on four.

In the expository sections, interpreters account for forty-one percent of the borrowed words, with Tillich and Wieman supplying the rest. In the comparative and critical sections, where King speaks as a critic in his own voice, interpreters account for eighty percent. His criticism of Wieman, that the word God carries three incompatible meanings in the system, comes from James Alfred Martin Jr.’s Empirical Philosophies of Religion, page 105. His argument that Wieman leaves no eternal conserver of value comes from Charles Hartshorne and William Reese, Philosophers Speak of God, pages 404 to 405, including the parenthesis calling an eternally habitable earth an astronomical impossibility. His central interpretive claim about Tillich, that Tillich is an absolute monist, together with the Hegel and Plotinus comparison and the three supporting quotations in order, is Boozer’s page 61. King’s own contribution is the personalist verdict stacked on top and the two-paragraph synthesis at the end.

Three transmitted errors settle the question of method. Randall misquoted Tillich twice and Boozer once, and each error appears in King’s text with a citation to Tillich. A man reading Tillich cannot reproduce another man’s mistakes. The editors’ headnote traces the practice to King’s notecards, on which he copied an interpreter’s reading, the primary quotation the interpreter had selected, and the interpreter’s footnote, then dropped the interpreter.

For Arday, the plagiarism was also clear. Times Higher Education reporter Jack Grove completed an analysis in September 2025 and produced a sixty-three page dossier of text comparisons; the story was killed after Arday engaged Carter-Ruck, a London defamation firm. David Sanders of Purdue, who reviews misconduct cases for that publication, said there was no question of extensive overlap. In June 2026 someone unhappy with Cambridge’s handling passed material to Nathan Cofnas (b. 1990), the philosopher Cambridge severed ties with in 2024. Cofnas published in July. The Telegraph reported a statistical analysis identifying 188 sentences in Arday’s thesis identical or highly similar to Paula Zwozdiak-Myers’s 2009 Brunel thesis. Jonathan Bailey, who runs Plagiarism Today and who called the campaign a weaponization of plagiarism against a Black academic, reviewed the documents and concluded that the evidence warranted a full investigation. Two of Arday’s papers had already been corrected. Overlaps were also reported with a 2013 paper by April Douglass, Dennie Smith and Lana Smith and a 2007 paper by Semiyu Aderibigbe, Laura Colucci-Gray and Donald Gray.

Now the institutional responses.

Boston University appointed four scholars in 1991: Robert Neville, John Cartwright, Charley Hardwick and Ray Hart. They reported that there was no question King had plagiarized by appropriating material from sources not credited, mistakenly credited, or credited too far from the borrowed passage. They estimated about twenty percent of the paper contained direct quotes or altered passages without proper attribution, noted that the final chapter exhibits very little borrowing, called the comparative part an intelligent contribution to scholarship, said revoking a doctorate in such circumstances would be absurd and unheard of, and recommended a letter be filed with the library copy. Provost Jon Westling accepted it.

Liverpool John Moores said the citation issues were the result of honest and reasonable error and that the overlap with Zwozdiak-Myers fell well within the accepted range of compliance with the academic standards of the time. Cambridge, notified in 2023, declined to investigate on the ground that the work predated Arday’s appointment, though he had been hired on the strength of it. Cambridge opened its investigation in August 2026, after three years of press.

An institution presented with evidence against a person it has elevated finds a frame that absorbs the evidence without touching the credential. Boston University used the merit of the surviving chapter. Liverpool John Moores used the standards of the era. Harvard used duplicative language for Claudine Gay. In each case the finding was scoped to the smallest area that would satisfy the complaint.

Universities do not adjudicate plagiarism against a person whose elevation they are invested in until an outside party forces it, and when they do act, the finding is scoped to preserve the credential. Boston University acted after Stanford’s editors and the Wall Street Journal forced it. Harvard acted after the New York Post. Cambridge acted after The Telegraph and The Times. Predict forward: where no outside party applies pressure, there will be no finding. If someone can produce a case where a university opened and published a misconduct finding against a celebrated appointee absent external pressure, the claim is wrong.

Clayborne Carson’s project was authorized by Coretta Scott King and staffed by scholars sympathetic to King. It found the plagiarism, sat on it, misled a Washington Post reporter, was scooped by Frank Johnson in the Sunday Telegraph on December 3, 1989, then confirmed the findings publicly and published everything, borrowing by borrowing, in a University of California Press volume with the source text printed underneath. That is what an honest archive looks like, and it took the project years and outside pressure to get there.

Cofnas is hostile to Cambridge, litigating against it, and has said that discrediting a Black professor would discredit the university that hired him for promoting DEI. Ghent University suspended him on August 20, 2026.

Bailey, who wants nothing to do with Cofnas’s politics, checked and found the evidence sufficient to warrant investigation. Grove, a staff reporter with no ideological stake, reached the same place a year earlier and was silenced by lawyers.

Ralph Luker, working on the King Papers, asked whether Crozer’s white faculty had graded King against a lower standard than Morehouse had, since his average rose from a C-plus to an A-minus when he moved to a mostly white institution. That question was raised in 1991 by a sympathetic editor. It is the same question being asked about Cambridge in 2026.

Five differences.

King was twenty-two years dead when the finding surfaced. Arday was forty-one and employed. A finding published about a dead man is history. A finding published about a living man is a sanction, and here it was delivered by press and Substack.

King’s standing rested on Montgomery, Birmingham and the Southern Christian Leadership Conference. The doctorate was an ornament. Arday’s professorship rested on the doctorate. Take away the thesis and you take away the appointment.

Verification runs the opposite way in the two cases. Boozer’s 1952 dissertation is not freely available, and the public depends on Carson’s apparatus and takes it on trust; King’s own typescript sits behind a ProQuest paywall. Zwozdiak-Myers’s thesis sits in Brunel’s open repository, downloadable by anyone. The Arday charge is more checkable by ordinary readers than the King charge has been.

Arday faced a second and separate charge, that details of his biography in a forthcoming memoir do not hold up. Fabrication and plagiarism are different offenses with different evidence, and running them together, as much of the coverage did, is how a citation dispute became a referendum.

Finally, the King case produced documents. A committee report, a scholarly edition, and a Journal of American History roundtable in June 1991 where David Garrow, Luker, John Higham, Keith Miller and Bernice Johnson Reagon argued it out in print. The Arday case has produced a killed article, a leaked dossier, a petition with six hundred signatures, a suspension in Belgium, and a death.

The Boston University committee’s 1991 report has never been posted. Liverpool John Moores has never released a finding. Cambridge has never released one.

Posted in Jason Arday, Martin Luther King, Plagiarism | Comments Off on A Passage-Level Count of King’s Dissertation, and What It Says About the Arday Case