Nathan Cofnas’s case against Jason Arday contains two accusations that are easy to collapse into one.
The first concerns authorship. Cofnas alleges extensive unattributed copying in Arday’s 2015 doctoral dissertation and in some later work. Those allegations can be investigated passage by passage, and they stand or fall on that evidence.
The second accusation is larger. After reviewing an Arday paper with no signs of plagiarism, Cofnas concludes that there is “nothing resembling real scholarship”. He points to the absence of elementary statistics and describes the research as recording experiences of racism and supplying commentary on them.
That second claim needs a control group.
Suppose Arday’s surviving original papers fall far beneath the methodological standards ordinarily demanded of professors in his field. The case that something unusual requires explanation gets stronger. Perhaps Cambridge and his previous employers suspended their normal standards.
Now suppose thousands of sociologists and education researchers publish work methodologically similar to Arday’s, and ordinary peer review treats it as scholarship. Plagiarism remains a serious question of integrity, and it explains nothing about the methodological character of the original work. The question becomes why an entire field accepts that kind of evidence.
Those are different stories with different remedies.
So I ran a simple experiment Cofnas does not perform. Instead of asking how Arday’s work looks from outside the sociology of education, I asked how it looks beside the work that passed through the same journal, the same editors and the same disciplinary culture. Then I read the one paper Cofnas has cleared of copying, on the theory that a paper nobody disputes the authorship of is the best available test of the second charge.
The results pull in two directions, and the essay follows both.
On his 2022 journal article, Arday looks like an ordinary practitioner of a field whose characteristic epistemic problem, when it appears, has little to do with sample size or the absence of statistics. The problem is the distance some authors allow between what they observed and what they claim to know. Call it claim inflation.
On the uncontested paper, he looks worse than that, and worse in a second way that has nothing to do with theory.
Cofnas’s easiest argument is his weakest. Arday interviews eighteen people. He runs no statistical analysis. Therefore, the suggestion goes, the result barely resembles research.
Qualitative interviewing is a recognized method with its own literature on recruitment, interviewing, transcription, coding, reflexivity, theme development, triangulation and the handling of contrary evidence. The COREQ reporting framework lists 32 items for interview and focus-group studies, covering sampling, the circumstances of data collection, recording, the derivation of themes, respondent validation and the use of supporting quotations. The broader SRQR framework sets out 21 reporting standards designed to let readers and reviewers evaluate a qualitative study. Neither says a study becomes scholarship once somebody calculates a p-value.
This bears on Arday because his 2022 article in the British Journal of Sociology of Education, “‘More to prove and more to lose’: race, racism and precarious employment in higher education”, contains a recognizable qualitative design. He studied eighteen academic staff of color across ten universities. Participants completed questionnaires and took part in interviews and focus groups. The material was recorded and transcribed. He describes deductive thematic analysis informed by Critical Race Theory, the construction of a coding frame and the iterative development of themes. The special-issue editors called the combination of surveys, interviews and focus groups a “complex and compelling methodology”.
One can dislike the method, find the interpretation circular, or think the evidence incapable of supporting the conclusions. The absence of statistical analysis establishes none of that. The better question is whether the method disciplines the inference.
Arday’s own journal supplies an almost perfect first control. His article appeared in a 2022 special issue on academic precarity, so we can compare it with papers accepted by the same editors, for the same issue, on the same subject.
Arday had eighteen participants. Nerida Spina and colleagues interviewed nineteen precariously employed academics in Australian universities, analyzing the material through Foucauldian ideas of power and discourse combined with a life-course approach, and recruiting through their own networks and Twitter. Martin Myers interviewed twenty-one Black and minority ethnic academics on zero-hours contracts, using grounded theory and the concept of White habitus. Catherine Oliver and Amelia Morris combined eleven interviews with their own autoethnographic experience to examine friendship and conferences. Aline Courtois and Marie Sautier used twenty-two interviews to study precarious migrant researchers around Brexit.
Arday’s eighteen is unremarkable in his immediate environment. An argument that begins with the small number of interviews and the missing statistics disqualifies a substantial share of the issue his paper appeared in.
The pattern holds across the journal. BJSE publishes intensive studies built on four pupils, five academics, six teachers, eight admissions tutors. The journal asks for work that is theoretically informed, methodologically rigorous and reflexive, and says submissions receive editorial screening followed by anonymized refereeing by at least two referees. Small-N qualitative work is one of the things the journal exists to publish.
This creates a trap for any critic of the field. It would be easy to trawl BJSE for tiny samples, list them with a sneer, and declare sociology fraudulent. Several of the strongest papers I read have very small samples.
Laura Quick follows four pupils identified as low attainers, studying them across several years with interviews and classroom observation, and keeps her conclusions tethered to what those cases can show. Paul Horton and colleagues focus on the relationship between two fifth-grade boys, combine ethnographic observation with interviews of teachers and students, and hold the interpretation to the social relation they observed.
Louise Archer and colleagues offer the sharpest comparison. Their project runs to more than two hundred longitudinal interviews with twenty working-class young people and twenty-two parents over eleven years, and their sample contains young people who became educationally mobile alongside those who did not. Their language about the role of luck stays correspondingly cautious.
Sally Riordan provides another model. Her larger project involved 152 interviews across thirty English schools. Cultural capital was never the interview subject. It surfaced unprompted in thirty-eight interviews at fourteen schools, and she then investigated what practitioners meant by it and how that compared with the research literature. The direction of travel runs opposite to the usual architecture: theory, theory-derived question, theory-compatible testimony, theory confirmed.
Sara Lindberg spent a year inside an international boarding school with participant observation and thirty-eight interviews. Her Bourdieusian reading finds relationships between social position and attitudes toward bilingualism, and she resists a deterministic account, discussing an observed case of habitus transformation that complicates the expected pattern. Bourdieu can lose, or at least be forced to accommodate what does not fit.
Small samples do not generate the problem. Strong theory does not generate it either. Pierre Bourdieu (1930-2002) appears in the most restrained papers I examined and in the least restrained. The dividing line is inferential discipline.
Once we stop demanding statistics, Arday’s vulnerability comes into view. His 2022 study has a heavily loaded epistemic architecture. Participants are academic staff of color, recruited through convenience sampling and recommendation. The study concerns race and precarious employment. Critical Race Theory foregrounds race and racism in the analysis. One interview question asks participants what role race or racism played in their experience of precarious work. The material is then read through a framework built around the significance of racial structures.
Asking people about racism is a legitimate research act. The trouble starts when we forget what the resulting evidence establishes. If participants say racism affected their employment, the study has good evidence that participants attribute their experience to racism, and often detailed evidence about the events behind those judgments. The participants’ causal account is a different object from an independently established cause. Perhaps they are right in every case. The study still needs an additional evidentiary step to show that racism produced a particular employment outcome rather than that respondents experienced and interpreted it that way.
Arday crosses that line at intervals. His themes point toward discriminatory and exploitative cultures in “the Academy,” and the paper moves between accounts of participants’ experience and language about systemic or institutional racism.
Then something happens in his limitations section. Arday says the findings do not generalize. He says that interviewing permanent lecturers, union officials, senior managers, human resources staff and employment agencies would have provided triangulation. He warns that quantitative work would be needed to “better discern causality”.
That admission changes the diagnosis. Arday understands the distinction between qualitative testimony and causal inference. He states it. The same paper carries epistemic caution and epistemic inflation, and that combination tells us more about the field than incompetence would.
The phrase “theoretically permissive” needs a definition, or it becomes a polite way of saying we dislike someone’s politics. Here is the pattern I mean. An observation sits on one side. Six teachers report this. Twenty-six trainee teachers experienced that. Five academics have these careers. One student underwent this transformation. A much larger proposition sits on the other. Racism caused the outcome. Neoliberalism produces the condition. White habitus explains the behavior of people nobody interviewed. Cultural capital reproduces an institutional advantage. A social process is sufficient to generate a transformation.
Sometimes the design supplies the bridge. The authors observe behavior over time, compare people with different outcomes, consult institutional records, triangulate competing accounts, or encounter cases that cut against the favored reading. Sometimes the theory supplies it. The observation gets treated as an instance of the framework, and the framework converts the instance into evidence for a general cause.
Myers gives the cleanest example in Arday’s own special issue. He interviews twenty-one BME academics on zero-hours contracts, a population well placed to describe its own experience of precarious work, departmental treatment and relations with permanent colleagues. His explanatory ambitions extend to the White academics he did not interview. The article explores how White habitus emerges as a shared collective trait within departments, argues that individual White academics act collectively to manage the risks of precarity through individual and collective racisms, and discusses White groups acting to hold their dominance. Those propositions may be true. The interviews did not sample that group’s beliefs or motives, and theory carries the account across the gap.
Fuad Arif Fudiyartanto and Garth Stahl supply a Bourdieusian version. They study five academics in one English department at one Indonesian university, two with overseas training and three without, interviewing them about professional biography, career progression and pedagogy. An intensive study of five careers can produce real knowledge. The abstract nonetheless generalizes about Indonesian academics with overseas training being more open to pedagogical innovation and advancing faster than their homegrown colleagues. The authors then call for a larger dataset across institutions, countries and disciplines. The methods section understands the evidentiary boundary. The abstract crosses it.
Biörn Ivemark and Anna Ambrose go smaller. Their 2023 article uses a theoretically sampled case study of one working-class student whose educational aspirations changed sharply, and uses it to develop an account of habitus transformation. A single case can illuminate a great deal; clinical medicine, anthropology and history would all be poorer without intensive study of exceptional cases. Their abstract says the case sheds light on some of the “sufficient conditions” behind dispositional disjunctures, and the body identifies processes that can sever the connection between habitus and its original social space. Sufficiency is a strong concept. One selected biography can show that a sequence occurred in a life and can make a proposed process plausible. Establishing sufficiency concerns what follows whenever specified conditions obtain, and that requires something more.
Ian Cushing gives a version worth taking seriously because his method is conscientious. He interviews twenty-six racially minoritized trainee teachers, records and transcribes, develops themes, and invites all participants to engage with his emerging interpretation; twenty-one do. He documents people being told to change their accents and their ways of speaking, and interviews capture that in a way no national dataset could. Then the explanatory language travels. Language oppression becomes a key reason England fails to retain racially marginalized teachers. The study does not measure retention. It contains no comparison between those who stay and those who leave, and no design for weighing language treatment against workload, pay, school conditions, geography or career opportunity. Member checking can establish that he has represented his participants fairly. It cannot establish that their experiences produced a national retention pattern. Evidence for a cause is a different thing from evidence about how much of an outcome that cause produces.
Rachel Stenhouse and Nicola Ingram examine how one private boys’ school prepares pupils for Oxbridge. Admissions statistics cannot show what happens inside elite schools, and observation can reveal the cultivation of comportment, confidence, vocabulary and ease with elite institutions. Three teachers volunteered for the principal observations and interviews, after three pilot interviews, and one question put to them asked whether they thought the sessions advantaged applicants to elite universities. The article presents itself as showing how private-school pupils acquire advantage in Oxbridge applications. The design can show what these teachers do and what they believe they are cultivating. There are no matched applicants without the intervention and no admissions outcomes to compare. Asking insiders whether their program works differs from demonstrating that it works.
An example I found is Gail Markle’s 2024 paper, “The sociopolitical liberalization of young adults: transforming a dominated habitus.”
Markle takes up the familiar charge that college turns conservatives into liberals. She recruited twenty-four college-educated Americans on three criteria: raised conservative, holding at least a bachelor’s degree, now identifying as liberal. She interviewed them about how the change happened, and reconstructs how they understand their own transformation. That is a legitimate qualitative question with a legitimate answer.
Her abstract says the findings “refute narratives of professorial or institutional indoctrination.”
The design cannot carry that. Every person in the sample was selected because she underwent the outcome to be explained. There are no conservative graduates who stayed conservative, no comparison across campuses with different political climates, no measure of exposure to professors’ political messages, and no one whose politics moved the other way. The causal evidence consists of retrospective self-explanation, which is not a measurement of influence. Try testing whether smoking causes lung cancer by recruiting twenty-four smokers who stayed healthy and asking them why. You would learn a great deal about twenty-four lives.
One qualification belongs to Markle. If the indoctrination narrative is the universal claim that every conservative who liberalizes at college was converted by a professor, counterexamples falsify it. The politically live hypothesis is probabilistic: that college, or particular campus environments, raise the likelihood of ideological change. An outcome-selected sample cannot touch that.
What followed bears on the diagnosis. BJSE included the article in its Paper of the Year winning collection, chosen by the executive editors from the previous year’s articles. That does not mean the editors endorsed the causal sentence I am criticizing. It does make one explanation harder to sustain, namely that a weak paper slipped past inattentive reviewers.
Everything above concerns work whose authorship nobody disputes on the part of authors nobody has accused. Arday’s 2022 BJSE article sits in that comparison as an ordinary case. There is a better test available.
Cofnas ran Arday’s output through detection software and reports that one paper came back clean: “Same storm, different boats: the impact of COVID-19 on Black students and academic staff in UK and US higher education,” written with Christopher Jones and published in Higher Education in October 2022. Whatever this paper shows about inferential discipline, it shows about Arday writing under his own steam. That makes it the strongest available evidence on the second charge, in either direction.
Read it and one passage does more work than any example I collected from BJSE.
The paper’s third theme concerns precarious employment. A US staff member, participant eleven, describes losing her job and attributes the redundancy to the pandemic. She says the redundancy “was because of the pandemic,” and adds that she noticed a lot of Black people and people of color being made redundant at her institution. The authors then write that “participant blamed the pandemic for their adverse circumstances but are absent in blaming their racist institution,” and offer this reading: “an interpretation could be the camouflage of institutional racism and how it can use the pandemic as an excuse to conduct racist practices by threatening their job security.”
Follow what happened to the evidence. A participant supplied a causal account that pointed somewhere other than the framework. The framework absorbed it. Her failure to blame institutional racism became evidence of institutional racism’s capacity to hide. On the test I proposed earlier, ask what observation would have counted against the favored explanation here. Testimony affirming racism confirms it. Testimony attributing the outcome elsewhere confirms it as camouflage. No answer a participant could give would register as a loss.
That is a stronger example than Myers or Cushing, where theory bridges a gap the evidence leaves open. Here theory converts contrary evidence into support.
The paper has more of the same architecture. Its first sentence asserts the permanence of systemic racism as established background. The design section says a CRT framework “guide[d] the structure of the interview schedule,” so the themes the analysis reports were the themes the instrument was built to surface. The conclusion then states that COVID-19 exacerbated all forms of racism and anti-Black racism in particular, a comparative and causal claim about a change over time that forty-odd interviews cannot measure. The limitations section says the findings do not generalize. The conclusion generalizes anyway, and does so about two national systems.
On the design-and-claim question, this paper sits at the top of my scale.
The paper also contains errors that have nothing to do with theory, politics or inference. They belong in their own category, because a defender can rebut a claim about inferential fashion by appealing to disciplinary convention, and cannot do that with arithmetic.
The sample is described three incompatible ways across two pages.
The participants section reports 43 participants: 18 staff and 25 students. Of the staff, it says 14 (78%) were female and 4 (32%) were male. Four of eighteen is 22 percent, and the two percentages sum to 110.
The next page describes the same sample as “eighteen students and twenty-four members of staff,” reversing the ratio and totalling 42.
The same paragraph then gives 26 Black women (16 students, 10 staff) and 16 Black men (7 students, 9 staff). That yields 23 students and 19 staff.
Three descriptions, none reconcilable with the others, of the sample on which every finding rests. A reader cannot say how many people were in this study or what proportion were staff.
The timeline does not fit the object either. Participants “were interviewed on multiple occasions between 2019 and 2021,” across what the paper calls a three-year period, about their experience of a pandemic that reached the UK and US in early 2020 and about a killing that occurred in May 2020.
The method is characterized as two things that exclude each other. The design section says CRT drove the analytic process and shaped the interview schedule. The same section invokes grounded theory, and describes grounded theory as rejecting positivist paradigms and deductive approaches. The paper wants the authority of emergence for themes that a stated framework was built to produce. Arday’s 2022 BJSE article, by contrast, calls its analysis deductive thematic analysis informed by CRT, so the tension runs between the two papers as well as inside this one.
The stated control for researcher bias is not a control. To limit their biases, “the researchers kept their subjectivity to not conform to one truth supporting different perceptions.” Whatever that sentence describes, no reader could apply it, replicate it, or determine whether it was done.
The citations do not hold up under checking. The Black tax is credited to Harper 1985 in the text and Harper 1975 in the reference list. Whiteness as property is credited to Harris 1985, which appears nowhere among the Harris entries, listed as 1993 and 1995. The same block quotation about the emotional toll of working harder is attributed to Bowden and Buie on one page and Bowden and Cullen on another, both at page 760. A footnote marker inside a quotation about the US president is numbered 5 where it should be 10. One reference entry begins “Public HealthBuikeme,” a collision of two sources. A sentence about Razai and colleagues lost its verb. Two of the empirical claims, on exploitation and on the abolition of precarious contracts, rest on a source given as “Arday, forthcoming,” which no reader can check, and the conclusion cites Arday 2022 to the same end.
None of these is fatal on its own. Together, in a paper offering three contradictory accounts of its own sample, they indicate a manuscript that nobody read closely, at the author’s desk or at the journal’s.
Higher Education is a Springer journal with peer review. The article is open access under a Creative Commons license and carries a declaration of no competing interests. It has been cited widely since publication.
I need to be careful here, because the temptation is to let a strong finding travel further than it should.
My thesis was that Arday looks ordinary within the qualitative sociology of education. On the 2022 BJSE article that holds. He sits in the large middle group, at 2 on my scale, alongside a dozen papers by scholars nobody has accused of anything.
On the uncontested paper it does not hold. That one scores 3, and the arithmetic and citation failures have no counterpart anywhere in my fixed sample. I screened three complete issues of BJSE. I did not screen an issue of Higher Education, so I cannot say whether its other articles show similar reporting failures, and I should not assume they do not. What I can say is that the sample contradictions in “Same storm, different boats” are of a kind I did not encounter in twenty-three BJSE papers.
So the honest formulation splits. Arday’s inferential habits are recognizably continuous with his field, and the strongest single instance of theory-proof reasoning I have found anywhere comes from him. His reporting standards in that paper fall below anything in my comparison set. The first observation is about a discipline. The second is about a paper, its authors and the journal that published it.
Keeping those separate protects the argument. Merging them lets a critic answer the weaker charge and claim to have answered both.
Here is a test anyone can apply before reading a paper’s findings. Read the theory and the method, and ask what result would weaken the preferred interpretation. If racism is the proposed cause, what evidence would count against it here? If White habitus explains the behavior, what observation would make the authors revise that account? If cultural capital explains elite advantage, what finding would favor a rival account? If neoliberalism explains precarity, what would count against it?
This asks for empirical constraint, not Popperian falsification from every ethnography. Qualitative methodology already contains the idea in negative-case analysis, where a developing explanation gets reworked in light of evidence that does not fit, and COREQ asks researchers to report respondent checking, the derivation of themes and the supporting evidence.
The best papers in my sample contain something capable of pushing back: longitudinal variation, contrasting participants, an unprompted finding, direct observation, contrary cases, multiple sources. The weakest do not. And the camouflage reading in “Same storm, different boats” shows the endpoint of a design with no such feature, where the only thing a contrary account can be is further proof.
Across these papers the authors appear to understand the problem. Arday knows his 2022 findings do not generalize and cannot establish causation, and says so in both papers. Fudiyartanto and Stahl know five academics cannot settle the larger question. Liuning Yang’s empirical object is his own autobiography, and his abstract moves from that autobiography to a claim about how the urban educational field constrains cultural capital among rural-to-urban migrant students, then calls for research on different subgroups.
So the inflation may live at the level of disciplinary rhetoric. A researcher collects evidence sufficient for a bounded claim. An article is expected to make a theoretical contribution. The author moves from the bounded finding to a statement about racism, neoliberalism, habitus or social reproduction. The limitations paragraph then retreats to what the study can support. The same paper says, in effect, that the research reveals a social process and that the research cannot establish causation. Peer review tolerates the tension because ambitious theoretical interpretation is one of the things the journal rewards. BJSE‘s own guidance asks for work well located within sociological theory while being methodologically rigorous.
The genre rewards theoretical generalization beyond the evidence. That proposition is testable sentence by sentence, and it is a more serious charge than incompetence.
This diagnosis is also not an outsider’s ambush. In 2021 BJSE published Tim Winzler’s critique of British Bourdieusian sociology of education, which argues that the tradition exhibits a distorted reflexivity and a poor handling of rival approaches and criticism. His argument is not mine, and it identified a tendency toward epistemic closure in this literature from inside.
I did not conduct the two-hundred-paper study needed to estimate how common this is. I began with something smaller and fixed the sample before judging it.
First, I compared Arday with the qualitative work published beside him in the 2022 special issue. Then I screened three complete ordinary issues of BJSE: 45(2), 44(5) and 45(6). Complete issues prevent me from hunting the journal for silly-looking articles. The issue determines the pool before any paper gets evaluated. My first-pass classification identified twenty-three qualitative or qualitative-dominant empirical papers across those issues. Mixed-method papers create a boundary problem, which is one reason this remains a pilot.
I coded claim stretch from 0 to 3. Zero means the central claim stays close to the people, setting and evidence studied. One means a broader interpretation offered as suggestion, possibility or theoretical application. Two means population, institutional, structural or causal claims that the design does not distinguish well from alternatives. Three means the paper claims to establish, explain or refute something its sampling or design cannot determine.
The distribution: four bounded, seven mild, nine moderate, three strong. Twelve of twenty-three carry a moderate or strong flag. Two are borderline; code both downward and it becomes ten of twenty-three.
Those numbers are not an estimate that half of the sociology of education is bad. There was one coder. The scale is not a validated instrument. I had already read some of the papers, so the coding was not blind. Full-text access varied. Classifying mixed-method work required judgment. Three issues are not a random sample of the journal, and the journal is not a random sample of the field.
What the pilot does establish is smaller and still useful. I did not have to search hard to find the phenomenon.
The scores, so you can attack them:
Coded 0: Louise Archer and colleagues on luck and educational mobility; Sally Riordan on the translation of cultural capital theory; Paul Horton and colleagues on bullying figurations; Gregor Schäfer and Katharina Walgenbach on educational strategies of upper-milieu German students, whose ninety-five interviews compare across milieus rather than sampling only the group whose behavior is to be explained.
Coded 1: Max Antony-Newman and colleagues on middle-class parental engagement; Bonita Cabiles on participation as relational investment; Andy Hamilton and colleagues on participatory action research with boys; Marta Cristina Azaola on Mexican technical schools; Jing Yu on Chinese international students and the U.S. racial hierarchy; Victoria de Leon Born and colleagues on autonomy and parental influence in educational choice; Sara Lindberg on bilingualism at an international boarding school.
Coded 2: Liuning Yang’s critical autoethnography; Gareth Burns and colleagues on working-class teachers; Munya Hwami and Michelle Bedeker on higher education in Kazakhstan; Saul Karnovsky and Brad Gobby on teacher wellbeing in a Reddit forum; Stenhouse and Ingram on private school entry to Oxbridge; Amy Stich and Andrew Crain on place-based habitus; Malin Ideland and Margareta Serder on affect in edu-business; Alireza Behtoui on empowerment and racialized segregation; Abdulaziz Aldossari on Saudi women’s choice of university majors.
Coded 3: Fudiyartanto and Stahl; Ian Cushing; Ivemark and Ambrose.
Three papers discussed above fall outside the fixed sample and should be treated separately, since I reached each one by a route the fixed-issue procedure exists to avoid. Arday’s 2022 BJSE article, from 43(4), I score 2. Markle’s, from 45(5), I score 3, and I found it by following the Paper of the Year collection. Arday and Jones in Higher Education I score 3, and I read it because Cofnas identified it as free of copying.
Disagree with any of these and the disagreement has to be about a specific inference in a specific paper. That is the point of publishing them.
In 1998 James Tooley and Doug Darby produced Educational Research: A Critique for Ofsted. Tooley examined 264 papers from four prominent education journals and analyzed forty-one in detail. Contemporary reporting said he found good practice in 31 percent, and he complained of small-scale, non-cumulative, poorly conceived projects.
Then came the counterattack. David Hustler and Ian Stronach went through Tooley’s work using his own criteria and accused him of inconsistency. Their most damaging point was that Tooley disclaimed generalization from his sample and then made sweeping statements about the health of educational research.
A critique of inferential overreach cannot rest on inferential overreach. My pilot is evidence that a phenomenon exists and deserves a larger audit. It is not evidence of prevalence.
One other feature of the surrounding field bears mention, independent of my argument. Matthew Makel and Jonathan Plucker examined the complete publication history of the hundred education journals with the highest five-year impact factors and found that 0.13 percent of articles were replications. A later mapping review covering 2011 through 2020 put the rate at about 0.20 percent, roughly one paper in five hundred. Much qualitative work is not designed for replication in the experimental sense, so this proves nothing about the papers above. It does show that concern about how education research checks its own claims predates the Arday affair.
Suppose a proper two-hundred-paper study eventually places Arday in the bottom 2 percent for inferential discipline. Cofnas’s argument gets stronger. We would then have evidence that Arday was doing something his field does not ordinarily accept, and one could ask why institutions rewarded an outlier.
Suppose instead he lands near the fortieth or fiftieth percentile. Diversity policy might still explain why he was hired, promoted quickly, celebrated or preferred over competitors; a methodological control group cannot settle personnel questions. It would become a poor explanation for why his research passed peer review. If ordinary scholars use comparable methods and make comparably expansive inferences, no diversity policy is needed to explain the journal’s acceptance of the work. The field was already built to recognize it as scholarship.
The uncontested paper complicates that fork in a way I did not anticipate, and the complication cuts both ways.
On the inferential question it supports the field-level reading. The camouflage move is a purer form of theory-proof reasoning than anything I found in BJSE, and it appears in an open-access article in a well-regarded Springer journal, where reviewers presumably saw the same sentences I did.
On the reporting question it supports Cofnas. Three irreconcilable descriptions of a sample, an interview window predating the events studied, and a bias control that says nothing are failures no disciplinary convention licenses. Nobody defends contradictory arithmetic as a house style. If work like this clears peer review and accumulates citations, either the reviewers did not read it or they read it and did not mind, and both possibilities need accounting for.
So the question shifts twice. First, why does the discipline treat the evidentiary move as sufficient? Second, how did a paper that cannot say how many people it studied pass review at all? The first indictment is less personal than Cofnas’s and much larger. The second is more specific than his, and it does not depend on plagiarism software.
That last point clarifies the role of the copying allegations. Plagiarism is serious misconduct and, if established, ends careers for good reason. It is orthogonal to the question here. Imagine two scholars producing equally weak papers, one copying passages and one writing every word himself. Plagiarism distinguishes their conduct. It does not distinguish the evidentiary quality of their conclusions. A detection program finds the copied sentence. It cannot find the missing inference, or the sample that adds to 110 percent. The paper Cofnas cleared is the one that fails hardest on both.
Treat what I have done as an exploratory audit. The initial question was whether Arday’s qualitative research looks unusually weak relative to research accepted in contemporary sociology of education. The first control group was the set of papers published alongside his 2022 article. The second was constructed by screening three complete issues rather than searching for examples. Papers were included when qualitative evidence formed a principal empirical basis of the article; purely quantitative and purely theoretical papers were excluded; mixed-method studies need a fixed inclusion rule in any replication. The three papers read outside the fixed sample are identified above along with the route by which I reached each.
The next study should use roughly two hundred qualitative articles. Specify the sampling frame in advance, across several years and at least two major journals. Freeze the rubric before coding. Strip author names, affiliations and explicit theoretical labels where feasible. Use at least two independent coders, report inter-rater agreement, and keep disagreements as data rather than reconciling them quietly. Do not identify Arday’s papers to the coders. Reveal theoretical frameworks only after the claim-stretch scores are complete, then classify papers as CRT, Bourdieusian, Foucauldian, otherwise theory-led, or comparatively theory-light.
That design lets several hypotheses compete. If Arday is an extreme outlier, the Arday-specific criticism survives. If CRT predicts greater claim stretch after matching on method and sample size, a CRT-specific criticism gains evidence. If Bourdieu, Foucault and CRT all behave alike, the problem belongs to theory-led qualitative inference in general. If theory-light papers perform the same, the problem is broader still. If none of it replicates, my diagnosis fails.
Alongside the 0-to-3 score, a replication should code separate yes-or-no variables: whether recruitment was described; what form the sampling took; whether the theoretical framework was specified before analysis; whether interview questions introduced the proposed explanation to participants; whether coding was described; whether more than one researcher coded or interpreted; whether evidence was triangulated against another source or population; whether negative or contradictory cases were reported; whether contrary testimony was reinterpreted as support for the framework; whether rival explanations were discussed; whether unsampled actors were assigned beliefs, motives or strategies; whether a participant’s causal attribution became an authorial causal assertion; whether the paper generalized from a local sample to a population or institution; whether causal language was used; whether the design contained anything capable of discriminating among plausible causes; whether the limitations section restricted generalizability; whether it disclaimed causal inference; and whether the abstract or conclusion claimed more than those limitations permit.
Two of those variables come out of the uncontested paper, and a replication should add a short reporting-integrity block alongside them, coded independently of any judgment about inference: whether the participant total is stated consistently everywhere it appears; whether reported percentages match the stated counts; whether the fieldwork period covers the events studied; whether the analytic approach is described consistently within the paper; whether cited sources appear in the reference list with matching dates and authors; and whether central empirical claims rest on unpublished work by the authors. These are mechanical checks. Two coders should agree on them almost always, which makes them a useful anchor for the harder judgments. COREQ and SRQR should be used to code reporting transparency, not converted into measures of truth.
Two further comparisons would strengthen or sink the argument. Code quantitative papers for the same failure, since a regression coefficient can be turned into a causal story as carelessly as an interview can, and qualitative sociology should not face a standard from which quantitative sociology is exempt. And compare abstracts and conclusions against limitations sections across a large corpus.
The investigation began with Jason Arday and is no longer mainly about him, though it has not left him behind either.
The easy story was that Cambridge elevated a man whose work bears no resemblance to what ordinary academics produce. On his 2022 journal article that story is hard to sustain. His methods look normal inside the journal that published him, and his inferential weaknesses look normal enough to have company. The field-level problem is weak calibration between research design and explanatory claim. In a substantial minority of qualitative papers, and possibly a large minority, theoretical frameworks supply causal, structural or institutional accounts that the underlying observations cannot distinguish from plausible alternatives.
The paper nobody disputes he wrote complicates the picture. It contains the clearest case of theory-proof reasoning I have found anywhere, and it also fails on checks that have nothing to do with theory or politics. Both findings need saying, and neither should be used to prove the other.
That formulation leaves prevalence open, and it lets the field defeat the criticism. A larger blinded sample may show my examples are freakish. Independent coders may reject my classifications. The effect may vanish in another journal. Bourdieusian, CRT and theory-light papers may show no difference. Quantitative work may show as much inflation in another form. A correction may yet appear that reconciles the sample figures in “Same storm, different boats,” and if one does, the strongest factual paragraph above goes with it. Those are empirical possibilities rather than rhetorical escape hatches. Constructing a control group means Arday is allowed to win, and so is sociology.
The question was never whether eighteen interviews are enough. Enough for what? If a study claims to describe what eighteen people experienced, eighteen may be plenty. If it claims to explain what caused their employment outcomes, fewer questions have been answered. If it claims to establish how an institution works, fewer still. And when a participant says the pandemic cost her the job, and the paper records her answer as proof that racism hides behind pandemics, the number of interviews has stopped being the issue. The evidence no longer has the power to tell the theory no.
