The Test That Was Not About the Test

In the early 1990s at Stanford, two psychologists handed undergraduates a set of very hard verbal questions from the Graduate Record Examination. Same questions for everyone. Same room. Same time limit. One sentence in the instructions changed.

Half were told the test measured their verbal ability. The other half were told the researchers were studying how people solve verbal problems, and that the test said nothing about how good they were at anything.

Claude Steele and Joshua Aronson reported the result. Black students who had been told the test measured ability scored lower than white students with matching prior SAT scores. Black students who had been told it measured nothing scored the same as those white students. Two further studies in the same paper showed why. Calling the test diagnostic switched on the racial stereotype in the students' minds, and it made them want not to be judged by it. In the fourth study of the set they published in 1995, simply asking students to record their race before starting was enough to move the scores [1].

They gave the effect a name. You have probably met it. Stereotype threat.

Here is what makes it interesting rather than merely sad. Nothing was done to those students. Nobody insulted them, nobody graded them unfairly, nobody told them a stereotype existed. It was already in the room, in the culture, in the air. All the experimenters did was make it relevant, and the scores moved.

That is a strange kind of cause. Nobody put it there. It is not inside the person and it is not really inside the test. It sits in the situation, and it only exists because everyone in the room already knows the same thing about everyone else.

The finding travelled fast. It is in undergraduate textbooks, in teacher training, in diversity workshops. Search for it now and the first page of Google is teaching centres explaining it to faculty, with a confident causal story and a list of things to do about it.

What that first page will not tell you is that the size of the effect has been argued about since 2004, that the largest studies keep coming back with nothing, and one of them with the opposite of the prediction, that the flagship classroom fix stopped working when somebody tried to repeat it properly, and that Steele and Aronson themselves never claimed the thing that gets attributed to them.

That argument is the article. Not because stereotype threat is fake, but because the true version is more useful than the confident one.

A Threat in the Air

Two years after the first paper, Steele wrote the theory out properly. The phrase he used in that 1997 paper has stuck: a threat in the air [2]. His claim was that stereotype threat is situational, not personal. It is not low self-esteem and it is not internalised belief in the stereotype. It is the pressure of knowing that a particular judgement is available to the people watching, and that your performance today could confirm it.

If that is right, the effect should show up for anyone, in any domain, as long as the stereotype exists and the person cares about doing well. That prediction was tested almost immediately, and it is where the theory earns its keep.

Steven Spencer, Claude Steele and Diane Quinn gave a difficult mathematics test to women who were strong at mathematics. Told the test had shown gender differences in the past, the women underperformed. Told it had shown none, they did not, and that 1999 result became the most cited application the theory has [3]. It is also the one that has been picked apart hardest, as we will see.

The same year, Joshua Aronson and four colleagues ran the version that should be better known. They took white men who were good at mathematics and told some of them that the study was about why Asian students outperform white students in mathematics. Those men scored lower in that 1999 experiment than white men who received no such framing [4].

Read that again. White men, on a maths test, underperforming because of a stereotype about somebody else.

Whatever stereotype threat is, it does not belong to minorities. It attaches to a specific comparison in a specific room, and it goes wherever that comparison goes.

Jeff Stone and his colleagues showed the same reversal on a golf course. Describe a putting task as a measure of natural athletic ability and white participants did worse. Describe the identical golf putting task as a measure of sports intelligence and Black participants did worse in the same 1999 study [5]. One task, two stereotypes, opposite victims depending on which one you invoked.

That symmetry is the best argument the theory has, and it is worth saying what it does and does not show. Not that the effect is large. That it is not about ability, because the same people move in whichever direction the comparison points. These are small laboratory studies, and later sections will be hard on studies that size. Publication bias selects on direction as well as on magnitude, so the symmetry is weaker evidence than it first looks. It is still the best the theory has, because it does not depend on the shift being large.

What Actually Happens in Your Head

Naming an effect is not explaining it. For years after 1995 the mechanism was a shrug: anxiety, presumably.

Then Toni Schmader and Michael Johns did something more specific. Across three experiments they measured working memory capacity directly, using a task that forces you to hold items in mind while doing something else. Priming the stereotype cut working memory capacity in women, and cut it again in Latino participants. In the third of those experiments, reported in 2003, the drop in working memory statistically accounted for the drop in maths performance [6].

That is a much better theory than anxiety, because it predicts where the damage will and will not fall.

Sian Beilock, Robert Rydell and Allen McConnell tested exactly that prediction. Stereotype threat hurt maths problems that lean on the phonological part of working memory, the part that holds verbal information while you manipulate it. It left alone the problems that did not. Better still, when participants practised the vulnerable problems until the answers came straight out of long-term memory rather than being computed on the spot, the effect went away, in the arm of their 2007 paper that tested exactly that [7].

The same person, the same room, the same stereotype, and the effect switches off because the arithmetic no longer needs the resource that was being occupied.

If you have read about cognitive load theory, this will sound familiar. Working memory is small, easily filled, and anything filling it competes with the task in front of you. Stereotype threat, on this account, is not a special psychological illness. It is a background process eating a resource you needed.

There is a physiological arm too. Jim Blascovich and colleagues measured blood pressure during an academic test in 2001 and found larger increases in mean arterial pressure among African American participants under threat, alongside worse performance on the harder items [8]. Jean-Claude Croizet and his team used heart rate variability as an index of mental workload, and found the load rose when a Raven's matrices test was described as measuring cognitive ability. Present the identical test as something else and the 2004 experiment found no gap between the two groups in that sample [9].

Bodies under evaluative pressure do this generally, which is the subject of what stress hormones do to memory. The stereotype does not create a new physiology. It supplies a reason for the ordinary one to switch on.

Toni Schmader, Michael Johns and Chad Forbes pulled the pieces into a single model. Three processes, running at once: a physiological stress response that directly degrades prefrontal function, a monitoring process that keeps checking your own performance for signs the stereotype is coming true, and the effort of suppressing the negative thoughts that monitoring throws up, in the model they set out in 2008 [10].

All three cost the same currency. That is the point of the model.

Yes

No

Cue in the room

Stereotype becomes relevant

Monitoring your own performance

Working memory occupied

Task needs working memory?

Score drops

No effect

Hold on to the diamond in that diagram. If the effect only appears when the task is hard enough to need the resource, then any real setting where tasks are calibrated differently should show a smaller effect, or none. Nobody had to wait for the sceptics to point that out. It falls out of the mechanism the theory's own supporters built.

The Sentence That Went Missing

Now the part that changes how you read everything above.

By the early 2000s, stereotype threat had escaped the laboratory and become the explanation for the racial gap in American test scores rather than one contributor to it. The version in circulation was that if you removed the threat, the gap would close.

In January 2004, Paul Sackett, Chaitra Hardison and Michael Cullen published a short paper in the American Psychologist pointing out that this is not what the 1995 study found [11]. Their own sentence is worth quoting exactly, because paraphrases have been the problem. The research is, they wrote, "widely misinterpreted in both popular and scholarly publications as showing that eliminating stereotype threat eliminates the African American-White difference in test performance".

The detail they pointed to is in the original paper's own methods. Steele and Aronson controlled statistically for prior SAT scores. That is reasonable when you want to isolate your manipulation. But it means the graphs show adjusted scores, not raw ones. Read the raw data and the picture changes: with stereotype threat removed, the two groups still differed by roughly the amount their earlier SAT scores predicted.

The manipulation moved performance. It did not close the gap. Steele and Aronson never claimed it did. It was never designed to test whether it could.

Steele and Aronson replied in the same 2004 issue, and their reply is not a retreat [12]. Their position is that no single experiment was ever meant to carry that weight, and that the accumulated literature, not the 1995 study alone, is what supports the broader claim. Sackett and his colleagues came back in 2005, still arguing that the field's summaries of its own findings had drifted from the findings themselves [13].

Both sides are being reasonable. Neither of them committed the misreading. Everyone downstream did, including a good deal of the first page of Google today.

This is a familiar shape. It happened to the bystander effect, whose founding story turned out to be wrong in its details, and to the Dunning-Kruger effect, whose original numbers reanalyse into something far less dramatic than the version people quote. A finding gets a memorable name, the name outruns the data, and correcting it afterwards is nobody's job.

The strongest counter-argument to Sackett is worth having on the page, and so is the objection to it. In 2009, Gregory Walton and Steven Spencer combined data from 18,976 students across five countries into two meta-analyses [14]. Their argument is that standard measures of academic performance systematically underestimate the ability of stereotyped students, because those measures are taken in exactly the environments that produce the threat. On that reading the ordinary test score is the biased number, not the adjusted one. The fair objection is that as stated the claim is hard to disconfirm, because any measured gap can be read as an underestimate of what is underneath it. Christine Logel and colleagues extended the argument to admissions policy in 2012 [15].

You do not have to settle that. You do have to know it is open.

Thirty Years in Fourteen Lines

Before the numbers, here is the shape of the argument, in order.

1995
Steele and Aronson name the effect across four studies
1997
Steele states the theory as a threat in the air
1999
Extended to women in mathematics and to white men
2001
Blood pressure rises during a test taken under threat
2003
Working memory identified as the mediator across three experiments
2004
Sackett and colleagues document the misreading and Steele replies
2006
A brief writing exercise cuts an achievement gap by 40 percent
2008
The first large meta-analysis puts the pooled effect at 0.26
2012
Only 30 percent of unconfounded experiments replicate
2013
Ganley and colleagues find nothing across 931 students
2017
Publication bias documented and the classroom fix fails replication
2018
A registered report on 2064 students finds no effect
2019
Under real testing conditions the effect falls to 0.14
2022
A raw-data meta-analysis of 31 studies finds no moderation

Notice how long the theory ran before anyone tried to replicate it at scale, and who eventually did. Sackett is a psychometrician, Wicherts and Flore are psychologists, and Borman and Hanselman are education researchers who had themselves published positive results. The corrections came from inside.

The Numbers, Side by Side

In 2008, Hannah-Hanh Nguyen and Ann Marie Ryan published the first large meta-analysis of stereotype threat experiments, and put the overall effect at d = 0.26, rising to d = 0.36 for women and d = 0.43 for minority participants on difficult tests [16]. That is about a quarter of a standard deviation overall. Oddly, for women the subtle cues did more damage than the blatant ones.

For eleven years that was the number people quoted, and it is not nothing. On a test scored out of 100 with a standard deviation of 15, it is about four points on your result.

Then Oren Shewach, Paul Sackett and Sander Quint asked a different question. Not how big the effect is in the studies we have run, but how big it is under the conditions that occur when someone sits a real high-stakes test.

Those are not the same conditions, and the difference is not subtle.

Bar chart of measured stereotype threat effect sizes from three meta-analyses

Data from Nguyen and Ryan 2008. Picho Rodriguez and Finnie 2013 and Shewach Sackett and Quint 2019. Chart by Mindomax.

Restricting the database to operational testing conditions in 2019 dropped the effect to d = 0.14, and adding the studies that used motivational incentives produced values ranging from 0.00 to 0.14 [17]. Real exams supply those incentives automatically. The same paper identified a previously unrecognised analytic error in studies that control for scores on a prior cognitive ability test, which had been biasing the estimate upward, and it found what the authors called nontrivial evidence of publication bias.

Their conclusion is careful and it is the sentence to remember. In operational settings such as college admissions and employment testing, the effect may range from negligible to small.

FeatureTypical laboratory studyA real high-stakes exam
How the threat is cuedTold directly that the test measures ability or shows group differencesNothing is said about groups at all
Test difficultySet near the ceiling on purpose so there is room to dropCalibrated across the whole ability range
What is at stakeCourse credit or a small paymentA university place or a job
Motivational incentivesUsually absentAlways present
Prior scores controlledUsually yesNever
Measured effectd = 0.26 pooled across all studiesd = 0.14 and some subsets 0.00

Look down the middle column and then the right one. Almost every feature a laboratory study needs to detect the effect is missing from an exam hall. That is no criticism of the laboratory work, which is supposed to strip away noise and amplify the signal. It is a criticism of reading a laboratory number as a field number.

Both numbers are answers. They answer different questions. Only one is about an exam hall, and it is the smaller one. The teaching-centre explainers that rank for this topic quote the first number, or no number at all. The second one has not reached them.

That gap matters most in the settings people care about, which is why it is worth reading alongside what actually happens to performance in high-stakes exams. The pressures in an examination hall are real. The question is whether this one is among the big ones.

What Happened When the Studies Got Bigger

The meta-analytic argument is about how to pool existing studies. A blunter question is what happens if you run one very large, carefully pre-planned study.

The answers have been consistent, and not comfortable.

Gijsbert Stoet and David Geary went back in 2012 to the study that started the women-and-mathematics literature and asked how often it had actually been replicated. Of the articles with designs capable of replicating the 1999 result, 55 percent did. Of those, half were confounded by the same statistical adjustment of prior mathematics scores that Sackett had flagged. Among the unconfounded experiments, their 2012 review put the replication rate at 30 percent [18].

Colleen Ganley and her colleagues ran three studies in 2013 with 931 students in total, using three different ways of activating the stereotype, from subtle to explicit [19]. Across all three, the mathematics performance of girls was unaffected. Gender differences turned up in two of the studies whether or not the stereotype had been activated at all.

Notice what is happening to the sample sizes. The original demonstrations were built on tens of participants. The studies finding nothing are built on hundreds, and the next one is built on thousands.

Then the biggest test of all. Paulette Flore, Joris Mulder and Jelte Wicherts pre-registered a study of 2064 students in Dutch high schools in 2018, committing in advance to the analysis and to the four moderators the field had proposed [20]. Those four were domain identification, gender identification, mathematics anxiety and test difficulty. They found no overall effect among the girls. They found no moderated effect either, on any of them.

A registered report is the strongest design here. It removes the freedom to hunt for the analysis that works.

Franca Agnoli and colleagues repeated a highly cited Italian study in 2021 with 328 students across ninth and eleventh grade, a much larger sample than the original, and reported a failure to replicate [21].

Every one of those nulls could still be explained away, and the standard explanation was that the theorised moderators had not been present in those particular samples. Somebody had to close that escape hatch, which meant testing the moderators themselves rather than the effect.

Andrea Stoevenbelt, Paulette Flore, Inga Schwabe and Jelte Wicherts did it in 2022, obtaining the raw data from 31 studies covering 3357 participants and asking whether the effect really is larger for people who identify with the domain and for whom the test is hardest [22]. It is not. They found no moderation.

That result costs the theory more than another null would. Test difficulty is not a side condition. It is the prediction the working memory account makes, the reason the diagram above has a diamond in it, and it did not hold up in the pooled raw data.

And then there is the study nobody predicted. James Chu and colleagues ran a field experiment in 2018 with 11624 students in Chinese vocational high schools, half of them primed about their educational track before taking technical and mathematics exams [23]. Vocational students in China are stereotyped as academically weak, so the prediction was straightforward. Priming had no effect on technical skills. It modestly improved mathematics performance.

Not no effect. A small effect in the direction nobody had predicted, on the largest sample you will find anywhere in this literature.

The authors' own explanation is that sorting students into a vocational track may crystallise the stereotype while removing academic performance as the thing they are measured on, which removes the threat. That is a post-hoc reading of an unexpected result, and it deserves the same discount as any other.

None of this makes the laboratory effect disappear. What it does is set a ceiling on one specific claim. Whatever a stereotype cue does to a school maths score, it is not large enough for a pre-registered study of two thousand teenagers to detect.

Seventeen Mediators

There is a second problem, about the shape of the theory rather than the size of the effect.

Charlotte Pennington, Derek Heim, Andrew Levy and Derek Larkin reviewed twenty years of work on what actually mediates stereotype threat. Their 2016 review found 45 experiments across 38 articles proposing 17 distinct mediators, sorted into affective, cognitive and motivational groups [24]. The strongest support went to anxiety, negative thinking and mind-wandering, all of which fit the working memory account.

Seventeen is a lot of mediators for one effect. Most theories manage with two.

The moderators multiplied the same way. Whether an effect appears is said to depend on how much you identify with the domain and the group, whether you endorse the stereotype, how hard the test is, whether you are the only one of your kind in the room, who is administering it, and how equal your country happens to be. Toni Schmader's 2002 finding that gender identification moderates the effect is a genuine result [25]. So is Michael Inzlicht and Talia Ben-Zeev's 2000 demonstration that simply putting a woman in a group with two men changed her performance [26].

But add them all together and you have a theory that can absorb any null result. The effect did not appear? The participants were not identified with the domain, or the test was not hard enough, or the cue was too blatant, or too subtle, or the country has a small gender gap.

Each is a testable claim. The trouble is that a theory with enough of them stops being falsifiable in practice, and the field noticed. Russell Warne put the sceptical case at its sharpest in 2021, responding to a meta-analysis that had reported an average effect of d = 0.28 among females [27]. His conclusion, in his own words, is that there is "no compelling evidence that stereotype threat is a real phenomenon in females".

That is stronger than the evidence supports, in the other direction. An operational estimate running from 0.00 to 0.14, in a literature with publication bias, is not an established zero, and Shewach and colleagues did not claim it was.

The honest position is narrower than either camp wants, and it fits in one sentence. Stereotype threat is a reproducible laboratory phenomenon whose real-world effect is small, and whose strongest claim, that it explains group gaps in test scores, is the part that has held up worst.

Someone Is Always on the Other Side of the Comparison

Almost every account of stereotype threat is written as something that happens to a disadvantaged group. That framing hides half the finding.

Gregory Walton and Geoffrey Cohen went back through the stereotype threat experiments and looked at the other participants. Not the group under threat. The group the comparison favours. Across the literature, those participants scored higher when the comparison was made salient than when it was not, an advantage they named stereotype lift in that 2003 review [28]. It pools the same experiments the later sections will question, so read it as the mirror image of a contested finding rather than as firmer ground.

The same room, the same instruction sentence, and it takes points from one group and hands them to another.

The most elegant demonstration of the two-sidedness came from Margaret Shih, Todd Pittinsky and Nalini Ambady. They tested Asian American women, who sit at the intersection of two stereotypes about mathematics that point in opposite directions. Prime the ethnic identity and performance went up. Prime the gender identity and it went down, in the 1999 experiment that made the point most cleanly [29].

One person. Two identities. Opposite results. It depended on which identity the situation made salient.

That study became a citation classic, and it is also a good illustration of how carefully this literature now has to be read. Carolyn Gibson, Joy Losee and Christine Vitiello attempted a replication with a much larger sample. Their 2014 report found the same pattern of means and a significant effect, but only after excluding the participants who did not know the relevant stereotypes existed [30].

That caveat cuts two ways, and it is worth being honest about both. You cannot be threatened by a stereotype you have never heard of, so excluding people who had not heard of it is what the theory predicts you should do. It is also a decision made after seeing the data, which is exactly the move this article complains about two sections later. A result that needs a filter is weaker than a result that does not.

The comparison is doing the work, which is the thread running through what social comparison does to performance. Stereotype threat is a special case of a general fact: telling people who they are measured against changes what they do.

The Ten-Minute Writing Exercise

If the effect is situational, you should be able to change the situation. That logic produced the most celebrated applied psychology of the 2000s, and then its most instructive failure.

The intervention is called values affirmation. Students spend ten or fifteen minutes at the start of a term writing about something that matters to them personally. Family, music, a friendship. Nothing about school, nothing about the stereotype. Reminding yourself what you value is meant to shore up your sense of being an adequate person, which makes the evaluative situation less threatening.

Geoffrey Cohen, Julio Garcia, Nancy Apfel and Allison Master ran two randomised field experiments in seventh-grade classrooms and published the result in Science in 2006. The writing exercise improved African American students' grades and cut the racial achievement gap in those classrooms by 40 percent across the two randomised field experiments [31].

Forty percent, from a fifteen-minute writing task, is the kind of number that gets a finding into every education policy document in the country. It did exactly that.

More followed. Akira Miyake and colleagues randomised 399 students in a college physics course in 2010 and found that values affirmation shifted women's most common grade from the C range to the B range, with the largest benefit to women who endorsed the stereotype [32]. Gregory Walton and Geoffrey Cohen tested a different intervention in 2011, a one-hour exercise framing social adversity in the first year of college as common and temporary, on 92 students of whom 49 were African American [33]. Over a three-year follow-up it raised the grade point averages of the African American students. Shannon Brady and colleagues followed the same participants into adult life and reported durable differences in 2020 [34].

Then somebody tried it at scale.

Geoffrey Borman, Jeffrey Grigg and Paul Hanselman ran the first district-wide implementation across 11 schools. The effects ran in the right direction, but their own 2016 summary is careful: the impacts were consistent with, and smaller than, those from the earlier small-scale studies [35].

The next result is the one that should have travelled and did not.

Paul Hanselman, Christopher Rozek, Jeffrey Grigg and Geoffrey Borman ran a well-powered replication in 2017, using the same procedures in the same setting where a previous large field experiment had produced significant benefits [36]. What they wrote is worth reading twice. "We found no evidence of effects in this replication study and estimates were precise enough to reject benefits larger than an effect size of 0.10." They then tested every moderator the theory offered, hunting for the condition that made the difference, and none of them explained it.

This is not a hostile replication by sceptics. Borman and Hanselman are two of the researchers who built the scale-up evidence in the first place.

And the story does not end there either. In 2018 the same group followed 920 students across the transition into high school and reported that self-affirmation reduced the growth of the racial achievement gap by 50 percent, with the largest effects in schools whose context cued stronger identity threat and among students who engaged more deeply with the writing [37].

The same team published a null and a large positive within a year. Both are honestly reported.

That is what a live scientific question looks like from the inside. Less satisfying than five evidence-based strategies, and much closer to the evidence.

The practical reading is not that affirmation exercises are worthless. It is that their effect depends on conditions nobody has pinned down, so an institution rolling one out district-wide should measure whether it worked rather than assume it did.

The Cheapest Fix Is Not the Student

One recent line of work has moved the question from the student to the room, and it is the most practically useful direction the field has taken lately.

Elizabeth Canning, Elise Ozier, Heidi Williams and Rashed AlRasheed manipulated something small in 2021: the beliefs a professor appeared to hold, as signalled in a course syllabus. In that experiment with 217 participants, both men and women read a fixed-mindset professor as more likely to endorse gender stereotypes, and both anticipated belonging less in the course [38]. Only for women did it translate into worse performance. The team then followed 884 real students across 46 STEM courses for two years and found the same pattern outside the laboratory.

Nothing about that intervention requires the student to do anything. The room changed. They did not.

It also lands close to something this cluster has covered before. What a teacher believes before meeting you can move your results, which is the Pygmalion effect. Two literatures separated for fifty years, converging on the same room from opposite sides. Pygmalion is the expectation somebody else holds. Stereotype threat is the cost of knowing it is available.

Older cue findings pointed the same way. Paul Davies and colleagues reported in 2002 that watching television commercials portraying women stereotypically depressed women's subsequent maths performance and steered them away from quantitative problems when given a choice [39]. That is media-priming work of exactly the vintage that has replicated worst, so hold it loosely.

The cue need not be about you. It has to be about the category you are in, and it has to arrive before the thing that gets measured.

The Field Changed the Subject

Since about 2020 the field has largely stopped fighting over mathematics test scores and moved into settings where the outcome is not a score at all. That move is easy to read as retreat, and it does follow the years in which the test-score claim came apart. Whatever the reason, it changes what the evidence can and cannot show.

Sarah Barber made the case in 2017 that this was necessary [40]. Age-based threat about cognitive decline, she argued, is not the same construct as the threat a minority student faces about intellectual ability. It is a threat to your self-concept rather than to your group's reputation, and the moderators that work elsewhere, group identification chief among them, do not generalise to it.

If she is right, treating stereotype threat as one thing with one set of rules has been an error the whole time.

The applied literature since then reads like a set of separate questions wearing the same name.

The most interesting of them measure a gap rather than a person. Gustav Tinghög and colleagues found in 2021 that measured gender differences in financial literacy shrank substantially once the threat was reduced [41]. If that holds, part of a difference everyone treats as real is an artefact of the measuring, which is the Walton and Spencer argument turning up in a field that was not looking for it.

Healthcare is now the busiest area, and the outcome measured there is health rather than a score. Adam Fingerhut and colleagues found in 2021 that lesbian, gay and bisexual patients who anticipated being stereotyped by clinicians reported worse health outcomes [42].

That is a real finding and a weaker design. It is self-reported and correlational, so it cannot separate the threat from everything else that comes with being treated badly by a health system. The laboratory work had the opposite problem: clean causation, trivial stakes.

Neuropsychological assessment is the sharpest case. Hannah VanLandingham and colleagues reviewed the literature in 2021 for clinicians who give cognitive tests to Black, Indigenous and other patients of colour, and pointed out that the test session carries every feature the laboratory studies use to create the threat [43]. A clinical decision may then rest on the score.

None of this tells you what it is like. For that you read the qualitative work, where the sentences belong to the people it is happening to.

Ebony McGee and Danny Martin interviewed 23 Black mathematics and engineering students in 2011 about how they kept succeeding in departments where they were constantly aware of being doubted [44]. One student's own words became the title of the paper. "You would not believe what I have to go through to prove my intellectual value."

Set that next to a meta-analytic effect of 0.14. The experiments measure what a stereotype does to a score in forty minutes. That student is describing what it does to a decade. Those are not the same quantity, and an argument about the first settles nothing about the second.

The motor-skill work keeps producing the reversal that made the 1999 studies persuasive, in domains where the stereotyped group is male. Priscila Cardozo and colleagues found in 2022 that a gender stereotype undermined balance performance in men [45], and Brenda de Pinho Bastos and colleagues found the same for dance in boys in 2023 [46]. Small studies again, with the same caveats, but the direction is the thing.

Whatever the size of the effect on a maths test, it was never confined to the groups the 1990s work made famous.

What Is Left Standing

So where does this leave you?

A few things can be said plainly, and they hang together rather than standing side by side. The phenomenon can be produced in a laboratory, and has been for thirty years, which is why it is situational rather than a trait: changing one sentence of instructions changes the result. Because it is situational it is also not confined to any particular group, which is what the white men against Asian men and the boys in dance are really showing. Where an effect does appear, working memory is the best supported of the seventeen proposed mediators, though the 2022 raw-data analysis weakened even that by failing to find the difficulty moderation the account predicts. And the number shrinks outside the laboratory either way, to somewhere between 0.00 and 0.14 under real testing conditions. Sitting underneath all of it is publication bias, which every reanalysis that went looking has found signs of, while disagreeing sharply about how much it moves the estimate.

The rest is genuinely open. Whether the effect is large enough to matter in a real exam hall. Whether it explains any meaningful share of group differences in achievement. Whether the moderators exist at all. Whether the classroom interventions work reliably, or only in conditions nobody has identified.

L. J. Zigerell's 2017 reanalysis of the Nguyen and Ryan database is the best model for how to hold all this at once. He found small-study effects, the signature of publication bias, and then applied four different correction methods. The four methods gave three different answers: essentially no change, a 50 percent reduction, and a reduction to near zero [47]. His conclusion cuts both ways on purpose, cautioning against citing the meta-analysis as evidence of a meaningful effect and against claiming the effect is negligible on the strength of these adjustments. Ann Marie Ryan and Hannah-Hanh Nguyen replied in the same journal in 2017, defending the original estimate [48].

Four methods, four answers, and an author willing to say so. That is more useful than a confident number.

There is one more thing worth saying, and it is not about effect sizes.

Suppose the operational effect really is 0.14, and suppose it explains almost none of the gap in test scores. The experiments would still have demonstrated something worth knowing. You can change how well a person performs on a hard task by changing one sentence about what the task means. Not by changing their ability, their preparation or their effort. By changing what they think is being measured, and who they think is watching.

That is a small effect with a large implication, and the large school studies put a hard limit on how far the implication reaches. They do not tell us which classrooms it applies to. They tell us that whatever it does in an ordinary school, the size of it is below what two thousand students can detect. What survives is the laboratory demonstration itself, which is smaller than the claim built on it and more interesting than nothing.

Frequently Asked Questions

What is stereotype threat in simple terms?

It is the drop in performance that can happen when you are doing something difficult in a situation where a negative stereotype about a group you belong to is relevant. Nobody has to say anything. Knowing the stereotype exists, and knowing your result could be read as evidence for it, is enough to occupy attention you needed for the task.

Who first discovered stereotype threat?

Claude Steele and Joshua Aronson named and demonstrated it in a 1995 paper in the Journal of Personality and Social Psychology. Steele set out the full theory in 1997, describing the effect as a threat in the air rather than something inside the person.

Is stereotype threat real, or did it fail to replicate?

Both, in a specific sense. The laboratory effect has been produced many times and working memory remains the best supported explanation for it, though a 2022 analysis of raw data from 31 studies failed to find the test-difficulty pattern that account predicts. But large pre-registered studies in real schools have repeatedly found nothing, including a registered report on 2064 Dutch students in 2018, and meta-analyses adjusted for publication bias give estimates that range from small to near zero. The strong claim, that it explains group gaps in test scores, is the part that has fared worst.

How big is the stereotype threat effect on a real exam?

The 2019 meta-analysis by Shewach, Sackett and Quint restricted the evidence to conditions that actually occur in high-stakes testing and found an effect of 0.14 standard deviations, with some subsets at zero. The commonly quoted figure of 0.26 comes from pooling laboratory studies, which are designed to make the effect appear.

What are examples of stereotype threat in a classroom?

Documented cues include being told a test measures ability, being told a test has shown gender differences before, being the only member of your group in the room, recording your race or gender before starting, and a syllabus signalling that the instructor believes intelligence is fixed. In each case the score changes without anything about the test changing.

Does stereotype threat only affect minority groups?

No. White men underperformed on a maths test when told the study was about why Asian students outperform them. White participants underperformed at putting when the task was framed as natural athletic ability. Men have shown it on balance tasks and boys in dance. It attaches to whichever comparison the situation makes salient.

What actually reduces stereotype threat?

The best-tested approach is values affirmation, a short writing exercise about something personally important. It produced a 40 percent reduction in a classroom achievement gap in 2006 and a 50 percent reduction in gap growth in a 2018 study of 920 students, but a well-powered replication in the same setting in 2017 found nothing and could rule out benefits above 0.10. Changing the situation rather than the student, such as how an instructor signals their beliefs, has more consistent recent support.

What is the difference between stereotype threat and stereotype lift?

Stereotype threat is the performance cost to the group the comparison disfavours. Stereotype lift is the matching gain to the group it favours, documented across the same body of experiments in 2003. The same instruction sentence produces both, which is why the effect is best understood as a property of the situation rather than of any group.