Philip N. Cohen has inspected the American Sociological Review three times to see whether its authors post the code and data behind their published results. The first inspection came in November 2020, the second in January 2024, the third in December 2025. Each time he read through the quantitative articles, looked for links to materials, and counted. He did not run any of the code. He says so.
The third post opens by reporting that the journal is making progress. Two-thirds of the quantitative articles in Volume 90 provided replication materials, he writes, against roughly a third in the previous audit and roughly a quarter in the first. Twelve of eighteen, up from eight of twenty-five.
The trouble is that the two numbers count different things.
In January 2024 Cohen sorted twenty-five quantitative articles into three columns. Eight had what he called a full package, by which he meant “data and code sufficient to replicate the papers.” Five or six others, depending on whether you read his prose or his table, supplied something less: partial data availability, a set of variable codes, fragments. The rest supplied nothing. The distinction between the first column and the second is the entire architecture of that audit.
In December 2025 the columns merge. His positive category now holds articles that “provide packages of code and data, or at least information about getting the data.” He adds that he counts materials parked on personal websites, which is not where replication files belong, and excludes materials described as available on request. Twelve of eighteen articles clear that bar.
I recoded the eighteen 2025 articles under the 2023 standard, one fixed rule, no code executed, matching his method. The rule: count an article when publicly accessible code appears to cover its quantitative results and the data are either included or obtainable by outside researchers through a documented procedure that does not run through the authors. Restricted data do not disqualify an article. A German scientific use file with a published application route counts. Swedish register data reachable through Statistics Sweden counts.
Under that rule, eight of the eighteen 2025 articles qualify. Forty-four percent, against thirty-two percent in 2023.
So the journal improved, but the improvement he reports runs from thirty-two to sixty-seven percent, a gap of thirty-five points, and the improvement a constant rule can find runs from thirty-two to forty-four, a gap of twelve. Roughly two-thirds of the reported rise comes from the widening of the category.
Four articles produce the difference, and none of them is bad work. Fabiana Silva, Irene Bloemraad and Kim Voss deposited materials for “Frame Backfire” at the Open Science Framework, and Cohen’s own table records that the project required an access request when he checked it on December 9, timestamped to the minute. He counted it anyway, and elsewhere in the same post tells the journal that permission-gated projects should not be allowed. Counting it gives nine of eighteen, which is the outer bound of the recount. Mabel Abraham, Tristan Botelho and James Carter posted data and code covering the main results of two of the studies in “(Not) Getting What You Deserve”; the article contains more than two. Daniel Scott Smith and coauthors, writing on “How Values and Uncertainty Shape Scientific Advance in Peer Review,” posted a repository whose own documentation says it holds no data, excludes some code, and cannot reproduce the machine learning benchmarks without material the authors supply on request. David Brady, Aliza Luft and Ezra Zuckerman Sivan, reanalyzing another team’s work in “How Does Culture Matter for Attainment, and How Would We Know If It Did?”, posted a PDF of their code, which states that the underlying data came from the scholars whose work they were reanalyzing and cannot be passed along.
In 2024 Cohen described his own audit as a tiny study of poor quality. In 2025 he calls his methods “squishy.”
Cohen has spent thirty years teaching other researchers not to do this.
The operation runs through his corpus from the beginning. In 2000 he and Lynne Casper built a new historical measure of cohabitation because the standard one misclassified people and distorted the trend it was supposed to reveal. In 2002 he showed that part of the apparent decline in the marriage premium for men came from cohabitors sitting inside the never-married comparison group, and that pulling them out shrank the decline. He and Lynne Casper split multigenerational households into hosts and guests because knowing that someone lives in an extended household tells you nothing about their position in it. He and his coauthors argued that a Black middle class defined through married households misses the growing population of never-married professionals living alone. He recoded keeping house as an occupation and found that women leaving unpaid domestic work accounted for as much of the decline in occupational segregation as desegregation within paid work. With Matt Huffman he took the established association between local Black population share and racial wage inequality and split it into two candidate processes, segregation and devaluation, and reported that one survived the data and the other did not.
A category that pools things behaving differently will manufacture a finding. Three cases show that he knows this. The first is PReF, the preprint review framework he built with collaborators, a fourteen-author project from 2022. It begins from the observation that peer review names procedures too different to compare, and replaces the label with eight standardized descriptors so a reader can see what a given review consisted of. The point of the exercise is to make comparison possible by first fixing what is being compared.
The second is his 2013 essay “The End of Men Is Not True: What Is Not and What Might Be on the Road Toward Gender Equality” on Hanna Rosin’s The End of Men, which is combative from the first page. The claim that young women out-earn young men turns out, on inspection, to describe single childless full-time metropolitan workers, and Cohen walks through the restriction before reanalyzing the narrowed group and finding that men still come out ahead within education categories. He does the same to her claim about managers, which rests on a category folding female-dominated professions into a statement about management. Polemic did not cost him the method. He goes claim by claim, category by category, for pages.
The third is closest to home. In 2023, with Micah Altman and Jessica Polka, he published “Interventions in Scholarly Communication: Design Lessons from Public Health” in First Monday. The complaint there is that open science reforms get launched without stable measures or systematic evaluation. The prescription is common measures, assessment designed in from the start, and comparable outcomes tracked over time. That appeared two years before the 2025 audit.
Applying his 2023 standard to his 2025 cases moves sixty-seven percent to forty-four.
Cohen has spent years asking sociologists to correct the record when the record is wrong, and has corrected his own audit twice when authors wrote to tell him he had missed their packages. I sent him the coding rule and the twenty-six cases, with links and dates, before publishing this. He declined to read it. The ledger stands open, and if any of it is wrong, it will be fixed here.
