Justin’s Notes

When the Bias-Hunters Leave Their Own Fingerprints

2026-07-21 · AI-generated insight

McIntyre's diagnosis of the social sciences includes a symptom he calls "ideological infection": "If one knows in advance what one wants to find, one will likely find it." He pairs it with Trivers's observation that when withheld data is finally examined, the errors run "most commonly in the direction of the researcher's hypothesis." Read Warne and Lewontin against that yardstick and something uncomfortable emerges.

Warne's account is the cleaner irony: Gould built a career exposing how Morton unconsciously fudged skull measurements to flatter his racial beliefs — and was then caught, by anthropologists, doing precisely that himself. The bias-hunter's own data ran in the direction of his hypothesis. This is McIntyre's rule turned on the very people who thought themselves immune, because their politics were the good ones.

Lewontin's famous paper is subtler, and more interesting, because you can watch the machinery operate in real time. He candidly admits his sample is skewed toward small "primitive" populations, that lumping decisions are "arbitrary," that reclassifying groups would raise the racial component — and then lands, reassuringly, on "6.3% is about right, or a slight overestimate." Every adjustment breaks toward the conclusion he wants. And the final sentence completes McIntyre's "questionable causation": a genetic partition statistic vaults, without argument, into a political prescription — "no justification can be offered for its continuance."

The pollination isn't that these scientists were wrong (Lewontin may be largely right). It's that the three notes together dissolve a comforting assumption: that anti-racist motives inoculate research against motivated reasoning. The measurement of prejudice is not exempt from the biases of measurement. Warne supplies the caught culprit; Lewontin supplies the live demonstration; McIntyre supplies the general law under which both fall.

Woven from these notes

Many scholars have criticized The Mismeasure of Man periodically throughout its 38-year history. For example, James T. Sanders stated that Gould’s attempt to link his argument to anti-racism was a ploy to smear intelligence scholars and Gould’s enemies as evil people. Arthur Jensen argued in 1982 that Gould misrepresented Jensen’s ideas and often demolished strawmen that no intelligence scholar believes, including the boogeyman of “biological determinism.” John Carroll showed that Gould understood neither the purpose nor interpretation of factor analysis (a statistical procedure often used to evaluate data from psychological tests) and that Gould’s attacks on factor analysis do nothing to alter the importance of intelligence tests, nor the mass of evidence—impossible to dispute—that they predict real-life outcomes. Most criticism of The Mismeasure of Man was confined to the recherché world of psychologists who study intelligence. However, a new debate opened up in 2011 when a team of anthropologists argued that Gould’s analysis of the data on cranium measurements from 19th century scientist Samuel George Morton was flawed. Gould cast Morton as a racist who fudged his data to match his beliefs about white racial superiority because of a supposed larger skull capacity. Instead, the anthropologists argued, it was Gould who manipulated the data to support his biases. This ignited a series of follow-up articles in the scholarly literature by authors taking a variety of positions regarding Morton’s data and Gould’s interpretations. Weisberg believed that the re-analysis was flawed and Gould was mostly correct. Kaplan and his colleagues claimed that Morton’s interpretations were flawed, but that Gould was incorrect in believing that he could discern Morton’s actions and motivations. Finally, Mitchell believed that Morton’s data were accurate and that the interpretations were colored by the racism of the era, but the claim that Morton subtly manipulated the data was a fiction created by Gould.
— Russell T. Warne The Mismeasurements of Stephen Jay Gould
When so many studies fail to be replicated, or draw different conclusions from the same set of facts, it does not instill confidence in the social sciences. Whether this is because of sloppy methodology, ideological infection, or other problems, the result is that even if there are right and wrong answers to many of our questions about human action, most social scientists are not yet in a position to find them. It is not that none of the work in social science is rigorous enough, but when policymakers (and sometimes even other researchers) are not sure which results are reliable, it drives down the status of the entire field. If medicine could break with its barbarous past, isn’t the same path open to the social sciences? For years, many have argued that if they could emulate the “scientific method” of the natural sciences, they too could become more scientific. But this simple advice faces several problems. Among the issues that plague contemporary social scientific research: Too much theory: A number of social scientific studies propose answers that have not been tested against evidence. The classic example here is neoclassical economics, where a number of simplifying assumptions — perfect rationality, perfect information — resulted in beautiful quantitative models that had little to do with actual human behavior. Lack of experimentation/data: Except for social psychology and the newly emerging field of behavioral economics, much of social science still does not rely on experimentation, even where it is possible. For example, it is sometimes offered as justification for putting sex offenders on a public database that doing so reduces the recidivism rate. This must be measured, though, against what the recidivism rate would have been absent the Sex Offender Registry Board (SORB), which is difficult to measure and has produced varying answers. This exacerbates the difficulty in (1), whereby favored theoretical explanations are accepted even when they have not been tested against any experimental evidence. Fuzzy concepts: Some social scientific studies can lead to misleading conclusions because of the use of “proxy” concepts for what one really wishes to measure. A recent example includes measuring “warmth” as a proxy for “trustworthiness,” in which researchers assumed — on the basis of studies which show that we are more likely to trust someone whom we perceive to be “on our side” — that perceptions of scientists as “cold” meant that they would be less trustworthy as well. But the two concepts may not be interchangeable. Ideological infection: This problem is rampant throughout the social sciences, especially on topics that are politically charged. Two ongoing examples are the bastardization of empirical work on the deterrence effect of capital punishment and the effectiveness of gun control on mitigating crime. If one knows in advance what one wants to find, one will likely find it. Cherry picking: The use of statistics allows multiple “degrees of freedom” to scientific researchers, but this is the most likely to be abused. In studies on immigration, for instance, a great deal of the difference between them is a result of alternative ways of counting the “costs” incurred by immigration. This is obviously also related to (4) above. If we know our conclusion, we may shop for the data to support it. Lack of data sharing: As the evolutionary biologist Robert Trivers reports in Psychology Today, there are numerous documented cases of researchers failing to share their data in psychological studies, despite a requirement from APA-sponsored journals to do so. When data were later analyzed, errors were found most commonly in the direction of the researcher’s hypothesis. Lack of replication: Psychology is undergoing a reproducibility crisis. One might validly argue that the initial finding that nearly two-thirds of psychology studies were irreproducible was overblown, but it is nonetheless shocking that most studies are not even attempted to be replicated. This can lead to difficulties, where errors can sneak through. Questionable causation: It is gospel in statistical research that “correlation does not equal causation,” yet some social scientific studies continue to highlight provocative results of questionable value. One recent sociological study, for instance, found that matriculating at a selective college was correlated with parental visitation at art museums, without explicitly suggesting that this was likely an artifact of parental income.
— Lee McIntyre To Fix the Social Sciences, Look to the “Dark Ages” of Medicine
The results are quite remarkable. The mean” proportion of the total species diversity that is contained within populations is 85.4%, with a maximum of 99.7% for the Xm gene, and a minimum of 63.6% for Duffy. Less than 15% of all human genetic diversity is accounted for by differences between human groups! Moreover, the difference between populations within a race accounts for an additional 8.3%, so that only 6.3% is accounted for by racial classification. This allocation of 85% of human genetic diversity to individual variation within populations is sensitive to the sample of populations considered. As we have several times pointed out, our sample is heavily weighted with “primitive” peoples with small populations, so that their Ho values count much too heavily compared with their proportion in the total human population. Scanning Table 3 we see that, more often than not, the Hpop values are lower for South Asian aborigines, Australian aborigines, Oceanians, and Amerinds than for the three large racial groups. Moreover, the total human diversity, Hspecies, is inflated because of the overweighting of these small groups, which tend to have gene frequencies that deviate from the large races. Thus the fraction of diversity within populations is doubly underestimated since the numerator of that fraction is underestimated and the denominator overestimated. When we consider the remaining diversity, not explained by within-population effects, the allocation to within-race and between-race effects is sensitive to our racial representations. On the one hand the over-representation of aborigines and Oceanians tends to give too much weight to diversity between races. On the other hand, the racial component is underestimated by certain arbitrary lumpings of divergent populations in one race. For example, if the Hindi and Urdu speaking peoples were separated out as a race, and if the Melanesian peoples of the South Asian seas were not lumped with the Oceanians, then the racial component of diversity would be increased. Of course, by assigning each population to separate races we would carry this procedure to the reductio ad absurdum. A post facto assignment, based on gene frequencies, would also increase the racial component, but if this were carried out objectively it would lump certain Africans with Lapps! Clearly, if we are to assess the meaning of racial classifications in genetic terms, we must concern ourselves with the usual racial divisions. All things considered, then, the 6.3% of human diversity assignable to race is about right, or a slight overestimate considering that Hpop is overestimated. It is clear that our perception of relatively large differences between human races and subgroups, as compared to the variation within these groups, is indeed a biased perception and that, based on randomly chosen genetic differences, human races and populations are remarkably similar to each other, with the largest part by far of human variation being accounted for by the differences between individuals. Human racial classification is of no social value and is positively destructive of social and human relations. Since such racial classification is now seen to be of virtually no genetic or taxonomic significance either, no justification can be offered for its continuance.
— Lewontin