What a Public Audit Does

In November 2020 Philip N. Cohen posted an inspection of the American Sociological Review. He had gone through the quantitative research articles in four issues of Volume 85, looked for links to replication materials, and found that four of fifteen provided them. He named the articles that did not.

In the same post he set out a theory of why researchers share. Authors share when they are personally compelled, when they are normatively compelled, or when they are formally compelled. The audit was an exercise in the first two. It made a public record of who had shared and who had not, on the premise that a scholar who sees his own paper listed under nothing will behave differently next time, and that a discipline that sees the list will develop an expectation.

He repeated the exercise in January 2024 and again in December 2025.

This note asks what those audits produced. It reports two tests of the personal-compulsion route, both of which return zero events in every arm. It reports a recount of the 2025 audit under the 2023 standard. It reports the policy history of the journal across the period, which rules out formal compulsion. And it reports a detection problem, discovered by accident, that bears on every audit of this kind including this one.

None of these is a large finding. Together they narrow the space of things a public audit can be doing.

The three audits do not use one rule. In January 2024 Cohen sorted twenty-five quantitative articles into three groups. Eight had what he called a full package, meaning data and code sufficient to replicate the paper, judged by inspection rather than execution. A further group supplied something less. The rest supplied nothing. His table and his prose disagree slightly on the middle group, giving five in one place and six in the other, with the residual moving correspondingly.

In December 2025 the middle group disappears. The positive category now holds articles providing packages of code and data, or information about getting the data. He notes that he counts materials hosted on personal websites and excludes materials described as available on request. Twelve of eighteen articles clear that bar.

He states the 2025 rule in the post. He then compares the resulting figure against the 2023 figure and reports that the journal is making progress.

I recoded the eighteen 2025 articles under the 2023 standard and recoded his eight 2023 positives under the same rule, so that the coder is constant as well as the criterion. The rule: count an article when publicly accessible code appears, without executing it, to cover the article’s quantitative results, and the data are either included or obtainable by outside researchers through a documented procedure administered independently of the authors. Restricted data do not disqualify an article; discretionary data do. A German scientific use file with a published application route counts. Swedish register data reachable through Statistics Sweden counts. Partial coverage of a multi-study article does not count. Materials not publicly accessible do not count.

All eight 2023 positives survive. Eight of the eighteen 2025 articles qualify.

That is 32.0 percent in 2023 and 44.4 percent in 2025, against a reported rise from 32.0 to 66.7. Counting the one article whose repository required an access request produces 50.0 percent as an upper bound. Roughly two thirds of the reported increase follows from the change in what was counted rather than from a change in what authors deposited.

The full case ledger, with the rule stated first, a resolving link and a check date on every entry, and the grounds for each downgrade named, is available separately. It was sent to Cohen before publication. He declined to review it.

The 2020 audit created an unusual natural comparison, and Cohen created it without meaning to. He inspected four of the six issues of Volume 85. The two he skipped contain quantitative research articles that were never listed, never named, and never assigned a sharing status in public. Assignment to the treated and untreated groups depends on which issues he happened to sample, which has no obvious relationship to any author’s propensity to deposit. Same journal, same volume, same editorial regime, same year, same author pool.

Volume 85 contains twenty-one quantitative research articles. Fifteen were audited. Eleven of those were baseline non-sharers named in public. Five comparable articles sat in the unaudited issues. One of those five had already deposited before the audit and is therefore not at risk. One further case, in the December issue, appeared close enough to the audit date that its exposure is ambiguous, and results are reported both including and excluding it.

The prediction, if public naming works through personal compulsion, is that named authors should subsequently deposit at a higher rate than unnamed ones.

The result is zero in both arms. No article in either group has a located first public deposit dated after November 30, 2020.

The second test uses the 2023 audit, where Cohen distinguished severity. Some articles were listed as supplying something partial. Others were listed as supplying nothing. Both groups were named in the same post. If naming operates through embarrassment, the second group had more to be embarrassed about. The result is again zero in both arms across the seventeen cases at risk.

This is not a null result on the treatment. It is the discovery that the outcome does not occur. Post-publication first deposit by ASR authors, across two cohorts and thirty-three articles, is close to a nonexistent behavior. Whatever a public audit accomplishes, it does not accomplish it by prompting authors to go back and post what they did not post before.

The design has no statistical power worth reporting and none is claimed. Eleven against four or five cannot detect an effect of any plausible size. What the comparison establishes is descriptive: in a hand-checked census of two full volumes, the behavior the intervention would have to change occurred zero times.

While cataloguing the corrections Cohen made to his audits, a pattern appeared that bears on all of this. The 2023 audit has been corrected four times in ways that changed an article’s classification. In each case an author of the article wrote in. In each case the material existed at the time of the audit and the method had not found it. Simone Zhang and David Johnson’s Harvard Dataverse deposit was published in January 2023, a year before the audit. Christof Brandtner’s materials had been available throughout. Ferry and colleagues’ OSF project predated the check.

A fourth case was found not by an author but by this recount. Cheng and colleagues were classified as supplying nothing. A Harvard Dataverse record associated with the article’s appendix had been public since August 11, 2022 and the article did not link to it.

Four detection failures in a volume of twenty-five articles is sixteen percent. Three surfaced because a particular author happened to complain. The fourth surfaced because someone re-ran the audit. There is no reason to think the four are all of them.

The implication runs in one direction and it runs at this note too. Audits that follow an article’s links measure what an article’s links reveal. My own searches used titles, DOIs and author names against six repositories, which is more than link-following, and Cheng shows the method sometimes recovers unlinked material. It gives no basis for estimating how often it misses. The eight of eighteen reported above is a floor. The true 2025 figure is probably higher, and the same holds for every figure in every audit of this kind. The comparison this note rests on is a comparison between two rules applied by one coder to the same articles, so a shared false-negative rate does not disturb it. Anyone reading the levels rather than the difference should read them as lower bounds.

The correction record is small and it has one clean feature. Across the three audits there are six documented corrections. Four changed a classification, and all four were prompted by an author of the paper being reclassified. One was self-initiated and arithmetic, catching an article counted twice. One came from a third party and fixed a broken link. No classification change originated with anyone other than an affected author.

The obvious reading is that authors are the people who notice, which is true and sufficient to explain the pattern. Six cases cannot distinguish that from anything else. What the record does establish is that the audit’s error-correction ran entirely through a channel available to people with standing to write to Cohen and an interest in the outcome. Cheng did not write in and the misclassification stood for two and a half years.

Cohen’s 2020 post argues that sociology operates on trust because almost nobody checks anybody’s work. His own audits were checked by the authors they audited and by no one else.

The remaining route in Cohen’s framework is formal compulsion, which would provide a straightforward institutional explanation for any rise between 2023 and 2025. It does not apply. The American Sociological Review did not adopt a mandatory deposit policy in the period. It does not require a data availability statement. It does not verify submitted materials. The American Sociological Association’s ethical standards ask members to make data available to qualified researchers after publication, which is a professional norm rather than a submission requirement. Sage operates a tiered research-data policy and leaves the tier to each journal; Acta Sociologica and Sociological Methods and Research adopted stronger tiers, and ASR did not.

The internal evidence agrees. Three of the eighteen quantitative articles in Volume 90 carry no data availability statement of any kind. That is what an unenforced regime looks like from the inside.

So across the window studied here, the journal imposed nothing, verified nothing, and the twelve-point rise that survives a constant rule occurred without a mandate. Two of the three routes Cohen named are ruled out or produced no measurable events. What remains is normative drift, or ordinary variation. Twelve points across denominators of eighteen and twenty-five is inside the range that sampling alone would produce.

Three questions that seem open are not. Whether deposit requirements raise deposit rates: they do. A recent large-scale coding of journal policies into tiers reports data availability at 16 percent where no policy applies, 87.5 percent where data are required, and 100 percent where a pre-publication reproducibility check is in force.

Whether verification raises reproduction: on present evidence it does not. In the same study, precise reproduction runs at 40.7 percent under no policy, 70.5 percent where data are required, 75.7 percent where data and code are required, and 65.0 percent under the strictest regime with a pre-publication check. Verification takes availability to complete and leaves a third of papers failing to reproduce precisely. These headline rates are conditional on obtaining usable data, which was possible for fewer than a quarter of the sampled papers, and should not be quoted without that clause.

Whether anticipation of inspection produces clean submissions: it does not. The Quarterly Journal of Political Science has reviewed submitted packages in house since 2005, on the stated theory that transparency would motivate authors to be cautious. Reviewing twenty-four papers, the journal found that twenty needed modification and fourteen produced results differing from the manuscript when the authors’ own code was run. The editor drew the conclusion himself: that these problems occurred despite authors knowing their code would be checked demonstrates the necessity of checking it. These are pre-publication states, caught and corrected, so the published articles are not defective; what the record shows is that anticipation was not sufficient.

The same shape appears elsewhere. The American Journal of Political Science, whose verification policy is the most studied in the social sciences, reported that eight of 127 submitted packages passed without revision. Sociology has one comparable case: for the Fragile Families Challenge special issue at Socius in 2019, two of fourteen manuscripts reproduced on the first attempt.

No study has publicly named individual authors for deficient sharing and then measured whether those authors subsequently deposited. The adjacent literature measures a different thing. Field experiments email authors privately and record whether materials arrive: roughly 38 percent in one large trial of over a thousand authors, 14 percent in another. Private request, private transfer, nothing public at either end.

The closest existing design is a randomized trial offering BMJ Open authors an Open Data Badge and measuring verified public deposit. Two of 54 in control, two of 57 in treatment. No effect. That tests positive recognition rather than public criticism, which are different treatments, so a null on one does not settle the other. It is the only randomized evidence on author-level reputational incentives and public deposit, and it points the same way as the zeros reported here.

The honest counterweight is that public naming demonstrably works on institutions. When a news organization singled out universities and hospitals for failing to report clinical trial results, reporting rates at the named institutions rose from an average of about 35 percent to about 76 percent; one university hired six compliance staff. Institutions have compliance offices, legal exposure and reputational management. Individual scholars have none of these, and whether the same lever moves them is the question nobody has answered.

The two comparisons here are underpowered by construction and are offered as bounded case series rather than hypothesis tests. The 2023 severity comparison carries a selection problem: articles in the partial group had already deposited something, articles in the nothing group had not, so the groups differ in the disposition the outcome measures and the confound runs against the prediction. Negative searches are the weakest form of evidence and the Cheng case demonstrates their failure mode. Where a deposit’s first public date could not be established from repository metadata or an archived capture, the case is recorded as unknown rather than resolved.

Repository states change. The pages that carry the most weight in the recount were submitted to the Internet Archive on August 26, 2026 and are cited by timestamped capture. One of them, an OSF project the audit had recorded as requiring an access request in December 2025, still redirected an anonymous visitor to a sign-in page eight months later. A GitHub repository named in the same audit stood at the same three commits. Nothing here can establish what any other repository looked like on any earlier date. The corrections catalogue rests on what Cohen published: post text, appended notes, and comment threads. Corrections made by email and never noted would not appear.

A public audit of a journal’s replication practices, repeated three times over five years by a prominent scholar, produced no located instance of an author going back to deposit. The journal it audited imposed no requirement and ran no check across the period. The rise the third audit reported is roughly two thirds an artifact of a widened category and, on a constant rule, is not distinguishable from ordinary variation. The audit’s own errors were corrected only by the authors it named.

None of that means the audits accomplished nothing. They may have shifted expectations in ways no design here could detect, and the discipline’s conversation about transparency is louder than it was. It does mean that the route Cohen named first, the individual scholar moved to act by seeing his paper listed, is not where the action is, and that six hours and a couple of hundred dollars per paper buys a check that a decade of counting has not.

The case ledger for the recount gives the coding rule, twenty-six entries, resolving links, check dates and archived captures.

About Luke Ford

I teach Alexander Technique in Beverly Hills (Alexander90210.com). Most of my posts since January 2025 are written with AI.
This entry was posted in Sociology. Bookmark the permalink.