Nathan Cofnas’s case against Jason Arday contains two accusations.
The first concerns authorship. Cofnas alleges extensive unattributed copying in Arday’s 2015 doctoral dissertation and in some later work. Those allegations can be investigated passage by passage, and they stand or fall on that evidence.
The evidence is what Cofnas says it is.
The second accusation is larger. After reviewing what he regards as Arday’s original work, Cofnas concludes that there is “essentially nothing resembling real scholarship”. He points to the absence of elementary statistics and describes the research as recording experiences of racism and supplying commentary on them.
That second claim needs a control group.
Suppose Arday’s surviving original papers fall far beneath the methodological standards ordinarily demanded of professors in his field. Perhaps Cambridge and his previous employers suspended their normal standards.
Now suppose thousands of sociologists and education researchers publish work methodologically similar to Arday’s, and ordinary peer review treats it as scholarship. Why would an entire field accepts that kind of evidence.
Those are different stories with different remedies.
So I ran a simple experiment. I asked how Arday’s work looks beside the work that passed through the same journal, the same editors and the same disciplinary culture.
Through this lens, Arday looks like an ordinary practitioner of a field whose characteristic epistemic problem is the distance some authors allow between what they observed and what they claim to know.
Call it claim inflation.
Cofnas’s easiest argument is his weakest. Arday interviews eighteen people (for “Same Storm, Different Boats: The Impact of COVID-19 on Black Students and Academic Staff in UK and US Higher Education,” Higher Education (2022)). He runs no statistical analysis. Therefore, the suggestion goes, the result barely resembles research.
Cofnas notes about the paper: “Copyleaks found no plagiarism.”
Qualitative interviewing is a recognized method with its own literature on recruitment, interviewing, transcription, coding, reflexivity, theme development, triangulation and the handling of contrary evidence. The COREQ reporting framework lists 32 items for interview and focus-group studies, covering sampling, the circumstances of data collection, recording, the derivation of themes, respondent validation and the use of supporting quotations. The broader SRQR framework sets out 21 reporting standards designed to let readers and reviewers evaluate a qualitative study.
This bears on Arday because his 2022 article in the British Journal of Sociology of Education, “‘More to prove and more to lose’: race, racism and precarious employment in higher education”, contains a recognizable qualitative design. He studied eighteen academic staff of color across ten universities. Participants completed questionnaires and took part in interviews and focus groups. The material was recorded and transcribed. He describes deductive thematic analysis informed by Critical Race Theory, the construction of a coding frame and the iterative development of themes. The special-issue editors called the combination of surveys, interviews and focus groups a “complex and compelling methodology”.
Does method disciplines the inference?
Arday’s own journal supplies a first control. His article appeared in a 2022 special issue on academic precarity, so we can compare it with papers accepted by the same editors, for the same issue, on the same subject.
Arday had eighteen participants. Nerida Spina and colleagues interviewed nineteen precariously employed academics in Australian universities, analyzing the material through Foucauldian ideas of power and discourse combined with a life-course approach, and recruiting through their own networks and Twitter. Martin Myers interviewed twenty-one Black and minority ethnic academics on zero-hours contracts, using grounded theory and the concept of White habitus. Catherine Oliver and Amelia Morris combined eleven interviews with their own autoethnographic experience to examine friendship and conferences. Aline Courtois and Marie Sautier used twenty-two interviews to study precarious migrant researchers around Brexit.
Arday’s eighteen is unremarkable in his immediate environment. An argument that begins with the small number of interviews and the missing statistics disqualifies a substantial share of the issue his paper appeared in.
The pattern holds across the journal. BJSE publishes intensive studies built on four pupils, five academics, six teachers, eight admissions tutors. The journal asks for work that is theoretically informed, methodologically rigorous and reflexive, and says submissions receive editorial screening followed by anonymized refereeing by at least two referees. Small-N qualitative work is one of the things the journal exists to publish.
This creates a trap for any critic of the field. It would be easy to trawl BJSE for tiny samples, list them with a sneer, and declare sociology fraudulent. Several of the strongest papers I read have very small samples.
Laura Quick follows four pupils identified as low attainers, studying them across several years with interviews and classroom observation, and keeps her conclusions tethered to what those cases can show. Paul Horton and colleagues focus on the relationship between two fifth-grade boys, combine ethnographic observation with interviews of teachers and students, and hold the interpretation to the social relation they observed.
Louise Archer and colleagues offer a comparison. Their project runs to more than two hundred longitudinal interviews with twenty working-class young people and twenty-two parents over eleven years, and their sample contains young people who became educationally mobile alongside those who did not. Their language about the role of luck stays correspondingly cautious.
Sally Riordan provides another model. Her larger project involved 152 interviews across thirty English schools. Cultural capital was never the interview subject. It surfaced unprompted in thirty-eight interviews at fourteen schools, and she then investigated what practitioners meant by it and how that compared with the research literature. The direction of travel runs opposite to the usual architecture: theory, theory-derived question, theory-compatible testimony, theory confirmed.
Sara Lindberg spent a year inside an international boarding school with participant observation and thirty-eight interviews. Her Bourdieusian reading finds relationships between social position and attitudes toward bilingualism, and she resists a deterministic account, discussing an observed case of habitus transformation that complicates the expected pattern.
The relevant dividing line here is inferential discipline.
Once we stop demanding statistics, Arday’s vulnerability comes into view. His 2022 study has a heavily loaded epistemic architecture. Participants are academic staff of color, recruited through convenience sampling and recommendation. The study concerns race and precarious employment. Critical Race Theory foregrounds race and racism in the analysis. One interview question asks participants what role race or racism played in their experience of precarious work. The material is then read through a framework built around the significance of racial structures.
Asking people about racism is a legitimate research act. The trouble starts when we forget what the resulting evidence establishes. If participants say racism affected their employment, the study has good evidence that participants attribute their experience to racism, and often detailed evidence about the events behind those judgments. The participants’ causal account is a different object from an independently established cause. Perhaps they are right in every case. The study still needs an additional evidentiary step to show that racism produced a particular employment outcome rather than that respondents experienced and interpreted it that way.
Arday crosses that line at intervals. His themes point toward discriminatory and exploitative cultures in “the Academy,” and the paper moves between accounts of participants’ experience and language about systemic or institutional racism.
Then something happens in his limitations section. Arday says the findings do not generalize. He says that interviewing permanent lecturers, union officials, senior managers, human resources staff and employment agencies would have provided triangulation. He warns that quantitative work would be needed to “better discern causality”.
That admission changes the diagnosis. Arday understands the distinction between qualitative testimony and causal inference. He states it. The same paper carries epistemic caution and epistemic inflation, and that combination tells us about the field.
The phrase “theoretically permissive” needs a definition. Here is the pattern I mean. An observation sits on one side. Six teachers report this. Twenty-six trainee teachers experienced that. Five academics have these careers. One student underwent this transformation. A much larger proposition sits on the other. Racism caused the outcome. Neoliberalism produces the condition. White habitus explains the behavior of people nobody interviewed. Cultural capital reproduces an institutional advantage. A social process is sufficient to generate a transformation.
Sometimes the design supplies the bridge. The authors observe behavior over time, compare people with different outcomes, consult institutional records, triangulate competing accounts, or encounter cases that cut against the favored reading. Sometimes the theory supplies it. The observation gets treated as an instance of the framework, and the framework converts the instance into evidence for a general cause.
Myers gives the cleanest example in Arday’s own special issue. He interviews twenty-one BME academics on zero-hours contracts, a population well placed to describe its own experience of precarious work, departmental treatment and relations with permanent colleagues. His explanatory ambitions extend to the White academics he did not interview. The article explores how White habitus emerges as a shared collective trait within departments, argues that individual White academics act collectively to manage the risks of precarity through individual and collective racisms, and discusses White groups acting to hold their dominance. Those propositions may be true. The interviews did not sample that group’s beliefs or motives, and theory carries the account across the gap.
Fuad Arif Fudiyartanto and Garth Stahl supply a Bourdieusian version. They study five academics in one English department at one Indonesian university, two with overseas training and three without, interviewing them about professional biography, career progression and pedagogy. An intensive study of five careers can produce real knowledge. The abstract nonetheless generalizes about Indonesian academics with overseas training being more open to pedagogical innovation and advancing faster than their homegrown colleagues. The authors then call for a larger dataset across institutions, countries and disciplines. The methods section understands the evidentiary boundary. The abstract crosses it.
Biörn Ivemark and Anna Ambrose go smaller. Their 2023 article uses a theoretically sampled case study of one working-class student whose educational aspirations changed sharply, and uses it to develop an account of habitus transformation. A single case can illuminate a great deal; clinical medicine, anthropology and history would all be poorer without intensive study of exceptional cases. Their abstract says the case sheds light on some of the “sufficient conditions” behind dispositional disjunctures, and the body identifies processes that can sever the connection between habitus and its original social space. Sufficiency is a strong concept. One selected biography can show that a sequence occurred in a life and can make a proposed process plausible. Establishing sufficiency concerns what follows whenever specified conditions obtain, and that requires something more.
Ian Cushing gives a version worth taking seriously because his method is conscientious. He interviews twenty-six racially minoritized trainee teachers, records and transcribes, develops themes, and invites all participants to engage with his emerging interpretation; twenty-one do. He documents people being told to change their accents and their ways of speaking, and interviews capture that in a way no national dataset could. Then the explanatory language travels. Language oppression becomes a key reason England fails to retain racially marginalized teachers. The study does not measure retention. It contains no comparison between those who stay and those who leave, and no design for weighing language treatment against workload, pay, school conditions, geography or career opportunity. Member checking can establish that he has represented his participants fairly. It cannot establish that their experiences produced a national retention pattern. Evidence for a cause is a different thing from evidence about how much of an outcome that cause produces.
Rachel Stenhouse and Nicola Ingram examine how one private boys’ school prepares pupils for Oxbridge. Admissions statistics cannot show what happens inside elite schools, and observation can reveal the cultivation of comportment, confidence, vocabulary and ease with elite institutions. Three teachers volunteered for the principal observations and interviews, after three pilot interviews, and one question put to them asked whether they thought the sessions advantaged applicants to elite universities. The article presents itself as showing how private-school pupils acquire advantage in Oxbridge applications. The design can show what these teachers do and what they believe they are cultivating. There are no matched applicants without the intervention and no admissions outcomes to compare. Asking insiders whether their program works differs from demonstrating that it works.
In Gail Markle’s 2024 paper, “The sociopolitical liberalization of young adults: transforming a dominated habitus,” she takes up the familiar charge that college turns conservatives into liberals. She recruited twenty-four college-educated Americans on three criteria: raised conservative, holding at least a bachelor’s degree, now identifying as liberal. She interviewed them about how the change happened, and reconstructs how they understand their own transformation. That is a legitimate qualitative question with a legitimate answer.
Her abstract says the findings “refute narratives of professorial or institutional indoctrination.”
The design cannot carry that. Every person in the sample was selected because she underwent the outcome to be explained. There are no conservative graduates who stayed conservative, no comparison across campuses with different political climates, no measure of exposure to professors’ political messages, and no one whose politics moved the other way. The causal evidence consists of retrospective self-explanation, which is not a measurement of influence. Try testing whether smoking causes lung cancer by recruiting twenty-four smokers who stayed healthy and asking them why. You would learn a great deal about twenty-four lives.
One qualification belongs to Markle. If the indoctrination narrative is the universal claim that every conservative who liberalizes at college was converted by a professor, counterexamples falsify it. The politically live hypothesis is probabilistic: that college, or particular campus environments, raise the likelihood of ideological change. An outcome-selected sample cannot touch that.
What followed matters to the diagnosis. BJSE included the article in its Paper of the Year winning collection, chosen by the executive editors from the previous year’s articles. It’s hard to argue that a weak paper slipped past inattentive reviewers.
The recurring pattern across these papers is that the authors appear to understand the problem. Arday knows his findings do not generalize and cannot establish causation. Fudiyartanto and Stahl know five academics cannot settle the larger question. Liuning Yang’s empirical object is his own autobiography, and his abstract moves from that autobiography to a claim about how the urban educational field constrains cultural capital among rural-to-urban migrant students, then calls for research on different subgroups.
So the inflation may live at the level of disciplinary rhetoric. A researcher collects evidence sufficient for a bounded claim. An article is expected to make a theoretical contribution. The author moves from the bounded finding to a statement about racism, neoliberalism, habitus or social reproduction. The limitations paragraph then retreats to what the study can support. The same paper says, in effect, that the research reveals a social process and that the research cannot establish causation. Peer review tolerates the tension because ambitious theoretical interpretation is one of the things the journal rewards. Its own guidance asks for work well located within sociological theory while being methodologically rigorous.
The genre rewards theoretical generalization beyond the evidence. That proposition is testable sentence by sentence.
In 2021 BJSE published Tim Winzler’s critique of British Bourdieusian sociology of education, which argues that the tradition exhibits a distorted reflexivity and a poor handling of rival approaches and criticism. His argument identified a tendency toward epistemic closure in this literature.
Here is a test anyone can apply before reading a paper’s findings. Read the theory and the method, and ask what result would weaken the preferred interpretation. If racism is the proposed cause, what evidence would count against it here? If White habitus explains the behavior, what observation would make the authors revise that account? If cultural capital explains elite advantage, what finding would favor a rival account? If neoliberalism explains precarity, what would count against it?
This asks for empirical constraint, not Popperian falsification from every ethnography. Qualitative methodology already contains the idea in negative-case analysis, where a developing explanation gets reworked in light of evidence that does not fit, and COREQ asks researchers to report respondent checking, the derivation of themes and the supporting evidence.
The best papers in my sample contain something capable of pushing back: longitudinal variation, contrasting participants, an unprompted finding, direct observation, contrary cases, multiple sources. The weakest do not.
I did not conduct the two-hundred-paper study needed to estimate how common this is. I began with something smaller and fixed the sample before judging it.
First, I compared Arday with the qualitative work published beside him in the 2022 special issue. Then I screened three complete ordinary issues of BJSE: 45(2), 44(5) and 45(6). Complete issues prevent me from hunting the journal for silly-looking articles. The issue determines the pool before any paper gets evaluated. My first-pass classification identified twenty-three qualitative or qualitative-dominant empirical papers across those issues. Mixed-method papers create a boundary problem, which is one reason this remains a pilot.
I coded claim stretch from 0 to 3. Zero means the central claim stays close to the people, setting and evidence studied. One means a broader interpretation offered as suggestion, possibility or theoretical application. Two means population, institutional, structural or causal claims that the design does not distinguish well from alternatives. Three means the paper claims to establish, explain or refute something its sampling or design cannot determine.
The distribution: four bounded, seven mild, nine moderate, three strong. Twelve of twenty-three carry a moderate or strong flag. Two are borderline; code both downward and it becomes ten of twenty-three.
Those numbers are not an estimate that half of the sociology of education is bad. There was one coder. The scale is not a validated instrument. I had already read some of the papers, so the coding was not blind. Full-text access varied. Classifying mixed-method work required judgment. Three issues are not a random sample of the journal, and the journal is not a random sample of the field.
What the pilot does establish is small. Arday did not emerge as an extreme case. On this rubric his 2022 article sits at 2, in the large middle group, with several papers in the comparison material easier to demonstrate as design-to-claim mismatches.
The scores, so you can attack them:
Coded 0: Louise Archer and colleagues on luck and educational mobility; Sally Riordan on the translation of cultural capital theory; Paul Horton and colleagues on bullying figurations; Gregor Schäfer and Katharina Walgenbach on educational strategies of upper-milieu German students, whose ninety-five interviews compare across milieus rather than sampling only the group whose behavior is to be explained.
Coded 1: Max Antony-Newman and colleagues on middle-class parental engagement; Bonita Cabiles on participation as relational investment; Andy Hamilton and colleagues on participatory action research with boys; Marta Cristina Azaola on Mexican technical schools; Jing Yu on Chinese international students and the U.S. racial hierarchy; Victoria de Leon Born and colleagues on autonomy and parental influence in educational choice; Sara Lindberg on bilingualism at an international boarding school.
Coded 2: Liuning Yang’s critical autoethnography; Gareth Burns and colleagues on working-class teachers; Munya Hwami and Michelle Bedeker on higher education in Kazakhstan; Saul Karnovsky and Brad Gobby on teacher wellbeing in a Reddit forum; Stenhouse and Ingram on private school entry to Oxbridge; Amy Stich and Andrew Crain on place-based habitus; Malin Ideland and Margareta Serder on affect in edu-business; Alireza Behtoui on empowerment and racialized segregation; Abdulaziz Aldossari on Saudi women’s choice of university majors.
Coded 3: Fudiyartanto and Stahl; Ian Cushing; Ivemark and Ambrose.
Two papers discussed above fall outside the fixed sample and should be treated separately. Arday’s 2022 article, from 43(4), I score 2. Markle’s, from 45(5), I score 3, and I found it by following the Paper of the Year collection.
Disagree with any of these and the disagreement has to be about a specific inference in a specific paper. That is the point of publishing them.
In 1998 James Tooley and Doug Darby produced Educational Research: A Critique for Ofsted. Tooley examined 264 papers from four prominent education journals and analyzed forty-one in detail. Contemporary reporting said he found good practice in 31 percent, and he complained of small-scale, non-cumulative, poorly conceived projects.
Then came the counterattack. David Hustler and Ian Stronach went through Tooley’s work using his own criteria and accused him of inconsistency. Their most damaging point was that Tooley disclaimed generalization from his sample and then made sweeping statements about the health of educational research.
My pilot is evidence that a phenomenon exists and deserves a larger audit.
One other feature of the surrounding field bears mention, independent of my argument. Matthew Makel and Jonathan Plucker examined the complete publication history of the hundred education journals with the highest five-year impact factors and found that 0.13 percent of articles were replications. A later mapping review covering 2011 through 2020 put the rate at about 0.20 percent, roughly one paper in five hundred. Much qualitative work is not designed for replication in the experimental sense, so this proves nothing about the papers above. It does show that concern about how education research checks its own claims predates the Arday affair.
Suppose a proper two-hundred-paper study eventually places Arday in the bottom 2 percent for inferential discipline. Cofnas’s argument gets stronger. We would then have evidence that Arday was doing something his field does not ordinarily accept, and one could ask why institutions rewarded an outlier.
Suppose instead he lands near the fortieth or fiftieth percentile. Diversity policy might still explain why he was hired, promoted quickly, celebrated or preferred over competitors; a methodological control group cannot settle personnel questions. It would become a poor explanation for why his research passed peer review. If ordinary scholars use comparable methods and make comparably expansive inferences, no diversity policy is needed to explain the journal’s acceptance of the work. The field was already built to recognize it as scholarship.
The question then stops being how Arday got away with it and becomes why the discipline treats this evidentiary move as sufficient.
The same reasoning clarifies the plagiarism allegations. Plagiarism is serious misconduct and, if established, ends careers for good reason. It is orthogonal to the field-level question here. Imagine two scholars producing equally weak papers, one copying passages and one writing every word himself. Plagiarism distinguishes their conduct. It does not distinguish the evidentiary quality of their conclusions. A detection program finds the copied sentence. It cannot find the missing inference.
Treat what I have done as an exploratory audit. The initial question was whether Arday’s qualitative research looks unusually weak relative to research accepted in contemporary sociology of education. The first control group was the set of papers published alongside his 2022 article. The second was constructed by screening three complete issues rather than searching for examples. Papers were included when qualitative evidence formed a principal empirical basis of the article; purely quantitative and purely theoretical papers were excluded; mixed-method studies need a fixed inclusion rule in any replication.
The next study should use roughly two hundred qualitative articles. Specify the sampling frame in advance, across several years and at least two major journals. Freeze the rubric before coding. Strip author names, affiliations and explicit theoretical labels where feasible. Use at least two independent coders, report inter-rater agreement, and keep disagreements as data rather than reconciling them quietly. Do not identify Arday’s papers to the coders. Reveal theoretical frameworks only after the claim-stretch scores are complete, then classify papers as CRT, Bourdieusian, Foucauldian, otherwise theory-led, or comparatively theory-light.
That design lets several hypotheses compete. If Arday is an extreme outlier, the Arday-specific criticism survives. If CRT predicts greater claim stretch after matching on method and sample size, a CRT-specific criticism gains evidence. If Bourdieu, Foucault and CRT all behave alike, the problem belongs to theory-led qualitative inference in general. If theory-light papers perform the same, the problem is broader still. If none of it replicates, my diagnosis fails.
Alongside the 0-to-3 score, a replication should code separate yes-or-no variables: whether recruitment was described; what form the sampling took; whether the theoretical framework was specified before analysis; whether interview questions introduced the proposed explanation to participants; whether coding was described; whether more than one researcher coded or interpreted; whether evidence was triangulated against another source or population; whether negative or contradictory cases were reported; whether rival explanations were discussed; whether unsampled actors were assigned beliefs, motives or strategies; whether a participant’s causal attribution became an authorial causal assertion; whether the paper generalized from a local sample to a population or institution; whether causal language was used; whether the design contained anything capable of discriminating among plausible causes; whether the limitations section restricted generalizability; whether it disclaimed causal inference; and whether the abstract or conclusion claimed more than those limitations permit. That last variable may prove the most productive of all. COREQ and SRQR should be used to code reporting transparency, not converted into measures of truth.
Two further comparisons would strengthen or sink the argument. Code quantitative papers for the same failure, since a regression coefficient can be turned into a causal story as carelessly as an interview can, and qualitative sociology should not face a standard from which quantitative sociology is exempt. And compare abstracts and conclusions against limitations sections across a large corpus.
The investigation began with Jason Arday and is no longer mainly about him. The easy story was that Cambridge elevated a man whose work bears no resemblance to what ordinary academics produce. The control group makes that story hard to sustain. His methods look normal inside the journal that published him.
The hypothesis that survives the controls is narrow. Arday does not appear unusually incompetent relative to the qualitative sociology of education published around him. The field-level problem is weak calibration between research design and explanatory claim. In a substantial minority of qualitative papers, and possibly a large minority, theoretical frameworks supply causal, structural or institutional accounts that the underlying observations cannot distinguish from plausible alternatives.
That formulation leaves the prevalence open, and it lets the field defeat the criticism. A larger blinded sample may show my examples are freakish. Independent coders may reject my classifications. The effect may vanish in another journal. Bourdieusian, CRT and theory-light papers may show no difference. Quantitative work may show as much inflation in another form. Those are empirical possibilities. Constructing a control group means Arday is allowed to win, and so is sociology.
The question was never whether eighteen interviews are enough. Enough for what? If a study claims to describe what eighteen people experienced, eighteen may be plenty. If it claims to explain what caused their employment outcomes, fewer questions have been answered. If it claims to establish how an institution works, fewer still. And when theory fills every gap between those propositions, the issue stops being the number of interviews. The issue is whether the evidence ever had the power to tell the theory no.
