{"id":200958,"date":"2026-08-19T17:00:12","date_gmt":"2026-08-20T01:00:12","guid":{"rendered":"https:\/\/lukeford.net\/blog\/?p=200958"},"modified":"2026-08-19T17:03:20","modified_gmt":"2026-08-20T01:03:20","slug":"what-the-arday-coverage-established-a-first-claim-level-audit-including-the-results-that-cut-against-my-own-thesis","status":"publish","type":"post","link":"https:\/\/lukeford.net\/blog\/?p=200958","title":{"rendered":"What the Arday Coverage Established: A First Claim-Level Audit, Including the Results That Cut Against My Own Thesis"},"content":{"rendered":"<p><a href=\"https:\/\/lukeford.net\/blog\/wp-content\/uploads\/2026\/08\/Jason_Arday_Claim_Audit_v2.xlsx\">Jason_Arday_Claim_Audit_v2<\/a><\/p>\n<p>NewsCord originally counted 249 substantive articles about Jason Arday (1985\u20132026) across fifteen national outlets. On August 19, 2026 it revised the figure to 229 for the period before he was found dead. The former count included nineteen written articles published after the Metropolitan Police were called at 3:12 p.m. on August 14, plus one video-only page. The workbook accompanying this essay preserves all 249 rows and flags the 229. Counting sounds objective until somebody asks where the boundary falls.<\/p>\n<p>Drawing the corpus at the moment of death frames it as what a man was subjected to while alive, and that framing is a causal premise sitting inside the sampling frame. No cause of death has been established and the case is with the coroner. A study that wants to describe press behavior rather than argue about its consequences should code the post-death window too, as a separate phase, and should say so before running any comparisons. The workbook now does the first half of that.<\/p>\n<p>NewsCord chose the outlets, the date range, and the threshold for substantive. Auditing a frame I did not build means inheriting its selection rules, and anyone replicating this should treat that as a limitation rather than a foundation.<\/p>\n<p>Now to what has been measured, and what remains promised.<\/p>\n<p>Three things are done. Every one of the 229 pre-death headlines has been coded for topic and for assertion mode, which is a census rather than a sample. Forty recurring propositions have been given an evidence status with a primary source and a stated caution. And twenty-eight atomic claim occurrences across nine articles have been fully coded on both scales, assertion strength and evidence strength as of the publication date, producing a calibration gap for each.<\/p>\n<p>Two hundred and twenty of the articles have coded headlines and unaudited bodies. Several are paywalled. The Guardian&#8217;s August 1 investigation carries amendments dated August 5, 6, and 11, which means the page available today is not the page a reader saw on August 1. Any statement about what the press told the public on a given morning requires archived publication-date versions, and until those exist the honest description of this project is a headline census plus an evidence ledger plus a nine-article pilot.<\/p>\n<p>That distinction governs everything below. Where a number comes from all 229, I say so. Where it comes from twenty-eight claims chosen partly for being checkable, I say that too, and the second kind of number cannot carry an argument.<\/p>\n<p>Begin with the census. Of the 229, thirty-four appeared before the resignation and 195 after. One hundred and five, close to half the corpus, appeared in the three days from August 5 to August 7. Whatever else the coverage was, it was a wave rather than an investigation, and the wave broke after the subject had already left his job.<\/p>\n<p>The first surprise concerns expansion. The narrative I had been developing, and which the commentary on both sides has adopted, says the inquiry began with scholarship and traveled outward into total biography until any statement Arday had ever made became newsworthy. At headline level that is not visible. Headlines whose subject is the personal biography, meaning publishing claims, threats and harassment, childhood, fundraising, and athletics, number twenty-two across the whole pre-death corpus. Six of the thirty-four before the resignation and sixteen of the 195 after. As a share, biographical headlines were more common in the early phase than the late one, seventeen percent against eight.<\/p>\n<p>What grew after August 5 was different. Institutional process and governance headlines went from three to fifty-one. Race, DEI, and culture-war framing went from two to twenty-four. Commentary went from five to thirty-six. Teaching and students went from none to eight. The post-resignation surge consists mostly of institutional consequence and opinion, which is what a story does when it stops being about a man and becomes about a university and a political argument.<\/p>\n<p>The second surprise concerns labels. Eighteen of the 229 headlines carry a categorical characterization by exact match against a fixed list: diversity poster boy, Professor Plagiarism, plagiarism professor or prof, fantasist, fabulist. Nine belong to the Daily Mail and MailOnline, seven to the Telegraph, one each to the Sun and the Express. Seven of the eighteen appeared before the resignation and eleven after, which against phase denominators of thirty-four and 195 means the labeling rate was roughly four times higher in the early phase. The Telegraph opened the mainstream cycle on July 24 with diversity poster boy and reused it. The characterization was set at the beginning and diffused. It did not escalate.<\/p>\n<p>Sixty-seven of the 229 headlines contain the stem plagiar, which is a better description of what the corpus was ostensibly about than any adjective.<\/p>\n<p>The version that says the press slid gradually from scholarship into a man&#8217;s childhood, growing more categorical as it went, is not what the headline census shows, and I had been about to publish that version. What the census supports is narrower: enormous volume, concentrated after the subject resigned, with the interpretive vocabulary established early by two outlets and repeated.<\/p>\n<p>Now the pilot, which produced a third result in the same direction.<\/p>\n<p>Across twenty-eight coded claim occurrences the calibration gap runs negative. Nineteen of twenty-eight show the article asserting less strongly than the available evidence permitted, six show a match, and three show overstatement. The mean is about minus eight tenths of a point. By publication: the Guardian sixteen claims at minus 0.81, the Times seven at minus 0.71, Times Higher Education three at minus 1.00, the Telegraph two at minus 1.00.<\/p>\n<p>Two reasons that number cannot travel. The nine articles were selected for being high-information and source-verifiable, which selects toward the careful investigative pieces and away from aggregation and commentary, and aggregation and commentary are where overstatement would live. And a sample of twenty-eight from a population of several thousand claim occurrences supports no estimate of anything. The finding is that in the corpus&#8217;s most substantial reporting, checked claim by claim against what was knowable on the day, the papers mostly claimed less than they could have. The Times article of August 5 is the clearest instance. It reported percentages summing to 110, inconsistent participant counts and identities, and a quotation attributed to different participants in different papers, and it then quoted an academic saying that whether this amounted to misconduct was for journals and universities to decide. A harsh story can be well calibrated.<\/p>\n<p>So the overstatement thesis survives only where it can be traced, and it can be traced in two places.<\/p>\n<p>The pig&#8217;s head is one. Arday said a severed pig&#8217;s head was delivered to his parents&#8217; house and that police had traced the animal through local butchers. The Guardian reported that the butcher he identified said police never visited and that the Metropolitan Police called the investigative account categorically incorrect. That contradicts the claim about the police investigation. It leaves the claim about the delivery unverified, since an absent police record is evidence against a report having been made rather than proof that nothing arrived. By August 12 the Guardian&#8217;s own wording had become that investigations found no evidence to support the pig&#8217;s-head claim, with the antecedent no longer specified. One proposition was disproved and a different one inherited the verdict. That migration happened inside a single outlet in eleven days and can be shown by quotation.<\/p>\n<p>The childhood story is the second. The Guardian found two men who remembered Arday speaking at primary school and two former school friends who supported his account, plus a woman who had helped him at three and described serious learning difficulties and no speech. The body reports the conflict. The headline leads with the challenge. There is no defensible coding of that evidence as false, and the correct status is disputed. What the article did is selection rather than error, which needs its own field in the codebook: contrary evidence known at publication, included or not.<\/p>\n<p>Those two instances make evidentiary contagion a hypothesis. Stated so it can lose: for claim families entering the corpus after a given date, first-mention assertion strength rises over the observation window while evidence strength for those same propositions stays flat or falls. The test needs body-level coding of all 229, first-mention identification per claim family, and a fixed rule for what counts as a family. If assertion strength turns out flat, the hypothesis fails, and given what the headline census and the pilot both show, failure is a live possibility.<\/p>\n<p>The evidence ledger meanwhile holds forty normalized propositions, and its distribution is the useful part. Eight are established, four strongly corroborated, two supported but unresolved, seven unverified, one genuinely disputed, nine contradicted, five substantially contradicted, and the remainder fall outside empirical adjudication: an interpretation, a causal claim about the death, a proposition with unresolved provenance.<\/p>\n<p>Cambridge&#8217;s February 2023 announcement is primary evidence that the university presented the appointment as a milestone in diversity, emphasized under-representation, and offered his career as inspiration. It is not evidence that racial diversity rather than ordinary academic considerations caused the hiring decision. Public relations language after an appointment is not a transcript of the electors&#8217; reasoning, and Cofnas&#8217;s stronger formulation belongs in the dataset as an interpretation. The same rule applies with equal force in the other direction. Paul Tiyambe Zeleza&#8217;s account of a man eagerly elevated as a symbol, ruthlessly scrutinized as an anomaly, and callously discarded as a liability contains an institutional description that overlaps with Cofnas&#8217;s to a degree that should embarrass both of their followings, and it also contains a causal claim about the death that no public evidence establishes and that Zeleza acknowledges is a matter for the coroner. Code the shared institutional sequence as an interpretation with substantial support. Code the causal claims as not adjudicable from public evidence.<\/p>\n<p>The Good Law Project petition belongs in the same ledger and gets the same treatment. Its assertion that investigations found no evidence whatsoever of wrongdoing is checkable against two published journal corrections and a documented sixty-three-page comparison dossier, and it fails. An institutional decision not to uphold misconduct establishes that a body declined to uphold misconduct. Liverpool John Moores considered allegations and did not uphold plagiarism, which is established and which the Guardian preserved correctly. Any headline treating that as a finding of guilt is wrong, and any petition treating it as proof of no underlying evidence is wrong in the mirror image. Both convert a contested record into a categorical verdict, and the study should score them identically.<\/p>\n<p>What the project now needs is controls, and three are missing.<\/p>\n<p>The first is a comparison corpus. Two hundred and twenty-nine articles in twenty-two days has no meaning without a baseline. Take a matched case, an accused academic at a comparable institution with a comparable evidentiary core, and count the same outlets over the same length of window. Claudine Gay (b. 1970) from the New York Post&#8217;s December 12, 2023 publication onward is one candidate, though the cross-national comparison introduces its own problems. A British case would be better. Without this, pile-on remains an impression with a large number attached.<\/p>\n<p>The second is a within-outlet label baseline. The Mail used a categorical label nine times. The question that makes the number mean something is how often the Mail applies categorical labels to any accused person of comparable prominence. If the answer is routinely, the finding is about the Mail. If the answer is rarely, the finding is about this story. The same test run on the Telegraph would settle whether diversity poster boy is a house habit or a case-specific frame.<\/p>\n<p>The third is reliability. Every number here comes from one coder who knew the hypotheses. Before any outlet-level comparison is published, a second coder should independently code a random subset of at least fifty claim occurrences, blind to the hypotheses, with agreement reported as a kappa and every disagreement published alongside its adjudication. The headline mode variable is the one I most expect to move, because the boundary between an event stated as fact and an opinion is where a coder&#8217;s priors do their work, and 155 of the 249 headlines sit in the mode I called opinion or general assertion, which is a suspiciously large residual category.<\/p>\n<p>Preregistration should come before any of that. Fix the codebook, the claim families, the phase boundaries, and the seven hypotheses, publish them, and only then obtain the archived article bodies. The resignation boundary is a cell in the workbook rather than a constant in a script, because sources give both August 4 and August 5, and a result that flips when the boundary moves one day is not a result.<\/p>\n<p>Where this leaves the argument.<\/p>\n<p>There was a substantial underlying story. Extensive unattributed textual overlap in the dissertation, two published corrections, internally impossible research data, disputed institutional affiliations, a book presented as published that was not, a six-day running claim the subject himself corrected to twelve, and a police investigation the police say never happened. Those are established or strongly corroborated on the ledger, and a press that ignored them would have failed.<\/p>\n<p>A second band remains unresolved. The thirty marathons, which Andy Murray was publicly describing in 2012 and which therefore was not invented for a Cambridge celebrity. The 600 miles over twelve days. The collective fundraising total. The fracture. The delivery of the pig&#8217;s head. The childhood speech. <\/p>\n<p>A third band consists of interpretations that no fact-check reaches.<\/p>\n<p>The results show a wave that arrived after the resignation, carrying institutional consequence and political argument, using categorical vocabulary that two outlets fixed in the first week. The claim-level migration I can prove happens inside single outlets on single propositions, twice, by quotation. <\/p>\n<p>The most useful thing a completed version could give either side is a measure nobody has: not whether an outlet was for or against him, but how fast each incorporated evidence pointing the other way. When Ohio State denied an affiliation, how quickly did it appear. When two schoolmates supported his account, how prominently was that carried by outlets running the two who did not. When a defender said there was no evidence whatsoever, did anyone place the journal corrections beside it. That test scores both camps on the same instrument, and it is the reason to publish the codebook, the ledger, the disagreements, and the adjudications, so that a reader who thinks Arday was a fraud, a reader who thinks he was destroyed by a racist campaign, and a reader who thinks both can download the same file and argue with the coding rather than with each other.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Jason_Arday_Claim_Audit_v2 NewsCord originally counted 249 substantive articles about Jason Arday (1985\u20132026) across fifteen national outlets. On August 19, 2026 it revised the figure to 229 for the period before he was found dead. The former count included nineteen written articles &hellip; <a href=\"https:\/\/lukeford.net\/blog\/?p=200958\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_monsterinsights_skip_tracking":false,"footnotes":""},"categories":[43259],"tags":[],"class_list":["post-200958","post","type-post","status-publish","format-standard","hentry","category-jason-arday"],"_links":{"self":[{"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=\/wp\/v2\/posts\/200958","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=200958"}],"version-history":[{"count":3,"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=\/wp\/v2\/posts\/200958\/revisions"}],"predecessor-version":[{"id":200962,"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=\/wp\/v2\/posts\/200958\/revisions\/200962"}],"wp:attachment":[{"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=200958"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=200958"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/lukeford.net\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=200958"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}