Rebecca Sear: ‘National IQ’ datasets do not provide accurate, unbiased orcomparable measures of cognitive ability worldwide’

Here is her paper, which is strongest where it is most boring, and weakest where it is loud.

Using Lynn and Becker’s own sample-type coding against them is the move that will survive. By their own generous definition, where 74 orphans in a single Eritrean orphanage counts as “national,” only 34 percent of the 656 samples qualify. That number comes from their spreadsheet. A defender has to attack his own data to attack it. The same goes for the case catalogue: Angola from 19 malaria-free individuals, Somalia from refugee children in Kenyan camps, Botswana from 140 adolescents recruited in South Africa on ethnic-proxy grounds. These are checkable in an afternoon.

Extending the catalogue to Denmark, Norway, Sweden, and Ireland closes the standard escape route. If Denmark’s number comes from fifth graders measured in 1968 and Norway’s from a cod liver oil trial and an epilepsy study, then the problem is the method.

Two arguments do most of the intellectual work. The first is non-independence. Twenty-five percent of current values are averaged from up to three neighboring countries, then fed into regressions that assume independent observations. That invalidates the published literature on statistical grounds alone. It is the argument an econometrician has to answer. She gives it a paragraph.

The second is the stability. The sub-Saharan average has stayed between 67 and 70 across four versions built from substantially different samples. Sample composition churns; output holds. That pattern is what a target value looks like. She gives it one sentence inside a longer passage about Wicherts. I would build the paper around it.

The fraud claim is where she overreaches. She makes it twice, in the introduction by way of Wicherts and in her own voice near the end. What the evidence supports is a research program with an announced prior, documented ad hoc adjustments toward that prior, and no correction after twenty years of criticism. Motivated construction. Fraud requires intent to deceive, and she has not established it to a standard that survives a hostile reading. The cost is practical. Say “fraudulent” and the defenders litigate the word instead of the sample sizes, and she is at a British university writing about a man whose institute still has friends.

She also gives the opposition its weakest case. Warne and Kirkegaard’s argument is convergent validity: national IQ correlates high with PISA, TIMSS, and adult literacy measures. A normal person looks at that and recognizes truth. Her answer is that the correlation tells us nothing. That is a weak argument. Random measurement error attenuates correlations toward zero. If the inputs were as noisy as she describes, the series should correlate with nothing. It correlates with a lot.

If comparable cross-cultural measurement is impossible in principle, then Wicherts’s corrected figure of 80 for sub-Saharan Africa is also meaningless, and so is the claim that Lynn biased the numbers downward, which presupposes a truer value. She wants the practical argument and the in-principle argument at once. They pull in opposite directions. The practical argument stands alone and I would let it. Measurement invariance is testable, and failures of it are findings rather than axioms.

Wicherts is doing heavy lifting for her as an ally while holding a position she rejects. His 80 is still low. His conclusion was that national IQs fail to support evolutionary explanations, and he thinks a regional estimate can be made and made better.

Slobodian (b. 1978) on Lynn’s influence is a claim about political history rather than about data quality. Including it tells you the target is journal editors and integrity officers rather than fence-sitting quantitative social scientists. That may be the right call given her goal, which is retraction. It makes the piece easier to file under politics.

The buried story is in the footnotes. Comparative Sociology published Jensen and Kirkegaard after the editor was told about the dataset. Nature Sustainability published Chu et al. in 2026 after Springer’s research integrity group was told. Two named journals, alerted in advance, publishing anyway. That is a gatekeeping story with dates and actors.

Small things. The DSM is the American Psychiatric Association, not the American Psychological Association. Warne appears as 2023 in text and 2022 in references. “a prior” for “a priori.” “the wholly inadequate of the primary data sources” in the conclusion is missing a word. She excludes Lynn for keeping control groups and for keeping treatment groups; both criticisms are fair, but stated together they read as scoring both ways.

The structural parallel to Jason Arday. Documented methodological failure, institution notified, institution proceeds, critic escalates from method to misconduct. The escalation is what happens when the first route fails.

About Luke Ford

I teach Alexander Technique in Beverly Hills (Alexander90210.com). Most of my posts since January 2025 are written with AI.
This entry was posted in IQ. Bookmark the permalink.