Introduction

Somewhere in your school career a stranger formed an opinion about you before you had said a word to them.

It might have come from a folder. A note from last year's teacher. A surname. An accent. A sibling who had been through the same room two years earlier and left an impression, good or bad. It might not even have come from the school. Parents hold expectations too, and a study following 9,654 adolescents through American high schools found that high parent expectations amplified the effect of high maths teacher expectations rather than competing with them [1]. Whatever the source, the belief arrived first and you arrived second, and for a while the person at the front of the room was reading everything you did through a lens you had no part in choosing.

That is the question this article is about, and it turns out to be a narrower question than the famous version of it.

The famous version says teachers' expectations shape how children turn out, and it comes with a story attached. In the mid 1960s two researchers walked into a California elementary school, gave the children an ordinary intelligence test, told the teachers it could identify who was about to bloom intellectually, named about a fifth of the children as bloomers entirely at random, and came back at the end of the year to find that the randomly chosen children had gained more than the others. The teachers had believed something false and the belief had made itself true.

That study is real. It happened. It is also the most attacked finding in the history of educational psychology, and the attacks landed. What survived is smaller than the story, stranger than the story, and considerably more useful.

The most interesting thing that survived is about timing. In 1984 Stephen Raudenbush pulled together eighteen experiments that had tried to induce teacher expectations and measure the effect on children's measured ability [2]. The thing that separated the studies where the effect appeared from the studies where it did not was not the age of the children, or the subject, or the country. It was how long the teacher had already known the child when the expectation was planted. The abstract puts it in one sentence: the better teachers know their pupils at the time of expectancy induction, the smaller the treatment effect.

Read that again with the title of this article next to it. What a teacher believes before meeting you is a different thing from what a teacher believes about you. The first can move you. The second is mostly just an accurate description of what you were already doing.

This article works through the whole argument. Where the popular telling is right, we will say so. Where it has been inflated into something more dramatic and less true, we will take it apart with the numbers. And where the field is genuinely still arguing, which it is in at least four places, we will name who is on which side rather than pretend the matter is closed.

One more thing before we start. Nothing here is advice about a specific child, and nothing here says a school can raise anyone's intelligence by trying harder to believe in them. This is an account of what has been measured.

Before the Children, There Were Rats

The Pygmalion effect did not start in a classroom. It started with a graduate student worrying about his own data.

Robert Rosenthal had a suspicion that experimenters were contaminating their own experiments. Not by fraud, but by wanting a result and unconsciously nudging things toward it. In 1963 he and Kermit Fode tested it in the least sentimental way available [3]. They gave student experimenters ordinary laboratory rats and a maze. Half the students were told their rats came from a strain bred to be maze-bright. The other half were told their rats were maze-dull. The rats were the same rats.

The maze-bright rats did better.

Nobody had lied to a rat. The rats could not read the label. Something the experimenters did with their hands, their patience, their willingness to record a run as a success or a failure, transmitted a belief to an animal that had no idea a belief was in the room. That paper appeared in Behavioral Science, volume 8, pages 183 to 189, and everything that follows is downstream of it.

The logical next step was obvious and uncomfortable. If a student handling a rat can do this, what does a teacher do to a child?

It is worth noticing how the question was framed from the very beginning. Rosenthal was not asking whether encouragement helps. He was asking whether a belief leaks. Those are different questions, and the second one is much harder to defend against, because you cannot decide to stop leaking.

Oak School

The classroom study was first published in 1966, in a short paper in Psychological Reports titled "Teachers' Expectancies: Determinants of Pupils' IQ Gains" [4]. Robert Rosenthal and Lenore Jacobson, who was the principal of the school, ran it across eighteen classrooms in a single elementary school in South San Francisco. Two years later it became the book everybody actually cites, Pygmalion in the Classroom, published by Holt, Rinehart and Winston.

The design was elegant, which is part of why it caught on.

Every child in the school took a test. The test was a real, ordinary measure of general ability. The teachers were told it was something else: an instrument that could identify which children were on the verge of an intellectual growth spurt. Roughly twenty percent of the children were then named to their teachers as likely bloomers, which in eighteen classrooms means about one child in five in every room. The names were drawn at random. Nothing about a named child's actual scores made them a bloomer. The label was the only thing that was real.

Note what the teachers were not told. They were never told a child was clever, or dull, or anything about current ability. They were told about a trajectory. That distinction matters later, because it is the reason the label could stick at all: a claim about the future cannot be immediately contradicted by what the teacher sees on Monday.

At the end of the school year the children were tested again. The children who had been named as bloomers had gained more than the children who had not, and the difference was largest in the youngest grades.

That is the study. Everything after this is an argument about what it means.

You can see why it travelled. It has a villain that is nobody's fault, a mechanism you can picture, and an implication that feels both scientific and kind. Findings shaped like that move faster than findings that are merely true.

Keep the design in mind as the criticism arrives, because most of the criticism is about measurement rather than about the idea.

The Review That Arrived the Same Year

Pygmalion in the Classroom was published in 1968. Robert L. Thorndike's review of it was published in 1968 as well, in the American Educational Research Journal, pages 708 to 711 [5]. The book and the demolition of the book landed in the same calendar year, and the demolition never got a fraction of the attention.

Thorndike's objection was not philosophical. It was arithmetic.

The instrument used to measure the children had been given to very young children in the lower grades, and in one class the average reasoning score came out in a range that would ordinarily describe intellectual disability. A whole class. This is not a subtle statistical point. It means the test was being read outside the range it was built to work in, and when a measure is behaving that badly at the bottom, the gains you compute from it are not measuring what you think.

The biggest reported gains in the study came from exactly those youngest grades. The place where the effect looked most dramatic was the place where the instrument was least trustworthy. That is not a coincidence you can wave away.

This is a general lesson about how big findings get built, and the 1960s produced several of them. The most quoted number in a paper is often the one from the smallest, noisiest corner of the data, because that is where the largest values live. The same decade gave us another famous social psychology result whose textbook version outran its evidence, and the pattern of how it happened is nearly identical. Noise is not symmetrical in its effect on a reputation. It only ever produces one headline.

This is why you should be careful with the number you have probably seen. Popular summaries of this study often report that first graders in the bloomer group gained over 27 IQ points. That figure is real in the sense that somebody computed it. It is also the single least defensible quantity in the whole literature, and it comes from precisely the grade where Thorndike showed the measurement was falling apart. If you read a page that leads with 27 points and does not mention Thorndike, you are reading a page that has not done the work.

A more defensible summary of the Oak School result, recomputed decades later in a review of the whole field, puts the overall effect on measured ability at about 0.35 standard deviations [6]. Real. Not nothing. Not 27 points.

My Fair Lady

The next thing that happens to a famous finding is that other people try it.

In 1971 Elyse Fleming and Ralph Anttonen published a large attempt in the American Educational Research Journal under the title "Teacher Expectancy or My Fair Lady" [7]. The title tells you what the temperature was. They did not find what Rosenthal and Jacobson had found.

They were not alone. Through the late 1960s and 1970s the replication record was ragged. Some studies found something. Many found nothing. Jean Elashoff and Richard Snow went back to the original data and published a book-length reanalysis in 1971, Pygmalion Reconsidered, brought out by Jones Publishing in Worthington, Ohio, and it did not flatter the original. Reviews of the field written at the time show researchers trying to work out whether they were looking at a real phenomenon or a statistical artefact with good publicity [8].

Notice what is happening to the argument at this point. Nobody is claiming the original result was fabricated. The claim is that it was fragile, and that a fragile result which everybody repeats stops being treated as fragile.

Five years later Charles West and Thomas Anderson raised the problem that would eventually turn out to be the deepest one in the whole area [9]. Their paper asked about preponderant causation. In plain language: when a teacher expects a lot of a child and the child does well, which way is the arrow pointing? Everyone had been assuming the expectation was doing the work. Nobody had established it.

Hold that thought. It comes back and it changes everything.

By the end of the 1970s Rosenthal and Donald Rubin had counted up 345 studies of interpersonal expectancy effects across every setting anyone had tried, not just classrooms [10]. The phenomenon was clearly not confined to schools. Whether it was large enough in schools to matter was a different question, and it was still open.

1963
Rosenthal and Fode show that experimenters get better maze performance from rats they believe are bred to be bright
1966
Rosenthal and Jacobson publish the Oak School result in Psychological Reports
1968
Pygmalion in the Classroom appears, and Thorndike's review takes apart its measurement in the same year
1971
Fleming and Anttonen fail to replicate, and Elashoff and Snow publish a book-length reanalysis of the original data
1976
West and Anderson ask which way the causal arrow actually points
1978
Rosenthal and Rubin count 345 studies of expectancy effects across every setting anyone had tried
1984
Raudenbush finds the moderator that explains the whole mess: how long the teacher already knew the child
1987
Wineburg and Rosenthal argue about it on consecutive pages of the same journal issue
1999
Spitz publishes a history of the controversy itself
2005
Jussim and Harber weigh thirty-five years of evidence and land on small, real and conditional
2025
Rubie-Davies and Hattie reopen the magnitude question from the observational side

So by 1980 the field had a famous finding, a pile of failed replications, and no agreement about what to do with either. What it needed was somebody willing to ask a different question.

The Number That Changed the Question

By the early 1980s the field had a pile of experiments that disagreed with each other. The usual move at that point is to average them and see what comes out. Stephen Raudenbush did something better. He averaged them and then asked what made them differ.

His 1984 synthesis in the Journal of Educational Psychology covered eighteen experiments on elementary school children, all of which had deliberately induced a teacher expectation and then measured children's ability, and it found a pooled effect of well under 0.2 standard deviations that varied enormously between studies [2]. Note the unit here, because it matters: eighteen studies, not eighteen children. The following year he and Anthony Bryk published the dataset again as the worked example in a paper on empirical Bayes methods for meta-analysis, and that version, with nineteen studies, is the one later statisticians have gone back to [11].

Here is what falls out of it. The numbers below come from a published reanalysis of that same nineteen-study dataset rather than from a summary of a summary, and the headline figure it reports is 0.078 standard deviations [12].

Pool all nineteen experiments together and ignore everything else about them, and that is what you get: an average effect on measured ability of 0.078, with a standard error of 0.052. A second estimation method gives 0.084 with the same standard error. Neither of those is statistically distinguishable from zero.

Stop there and you would conclude the whole thing was a mirage. Almost nobody does stop there, and they are right not to, because the studies vary among themselves far more than sampling noise can explain.

Now add one piece of information about each study: how many weeks the teachers had already spent with those children before the false expectation was planted. Suddenly the picture resolves. At the average amount of prior contact across the studies, the effect is 0.134 standard deviations, and it is statistically significant. And the slope on prior contact is negative and also significant: minus 0.157 for every additional week the teacher had already known the class.

That slope is the article. It is steep. A teacher who has spent a couple of extra weeks with a group of children before the label arrives is a teacher for whom the label does very little.

Effect on measured ability, in standard deviations, by what is being comparedPooled experimentsSame, prior contact modelledOak School restatedNine meta-analysesHigh vs low expectation teachers10.90.80.70.60.50.40.30.20.10Standard deviations

Those five bars are not five measurements of one thing, and reading them as though they were is the single most common mistake made about this topic. The first three come from experiments that planted a false belief in a teacher's head. The last comes from comparing different teachers with each other, which is a completely different kind of evidence and cannot be read causally at all. The fourth pools studies of both kinds. A later section is about why that distinction matters more than anything else on this page.

Two Weeks

Why would prior contact matter so much? Because an expectation is a hypothesis, and a hypothesis only survives where there is no data.

A teacher who has never met a child has nothing to weigh a claim against. Told that this one is about to bloom, they have no reason to disbelieve it and no evidence to contradict it. The belief sits there, uncontested, shaping how they read every ambiguous thing the child does for the next few weeks. And ambiguous things are most of what happens in a classroom.

You can test this against your own experience with new people. The first thing you are told about someone you have never met does an enormous amount of work. The fifth thing you are told about someone you already know does almost none.

A teacher who has had that child in front of them for a month is in a different position entirely. They have watched the child struggle with something and get it. They have seen what happens when the room gets noisy. They have marked twenty pieces of work. Now somebody hands them a note saying this child is about to bloom, and the note has to compete with a month of first-hand observation. It usually loses.

This is a specific instance of something more general about how impressions form. A first piece of information about a person does disproportionate work, because everything after it gets interpreted in its light rather than weighed independently. The same asymmetry shows up in how a first number distorts every estimate that follows it, and it is why a single global impression can drive judgements about unrelated qualities. Once a teacher holds a belief, ordinary confirmation bias keeps it alive: the child's good day is evidence, the child's bad day is an off day.

Modern work has gone at the same question directly. Anneke Timmermans and colleagues asked in 2021 whether teachers adjust their expectations as the year goes on or hold on to the impression they started with [13]. The title of the paper is the question: adjusting expectations, or maintaining first impressions? That the question is still worth a paper in 2021 tells you it is not settled.

There is a practical reading of all this that has nothing to do with psychology and everything to do with school administration. The moment of maximum vulnerability is the handover. The end-of-year note, the transfer file, the streaming decision made before the new teacher has met anyone. That is the window in which a belief about a child is doing the most work, precisely because nobody in the room has any evidence yet.

That is the mechanism at the level of evidence. Now the mechanism at the level of a room.

How a Belief Gets Out of a Teacher's Head

None of this works unless the belief travels. Nobody tells a seven-year-old they have been assigned to the low group. So how does the child find out?

Jere Brophy laid out the standard account in 1983 [14]. Teachers who expect more of a child treat that child differently, in ways that are individually small and collectively substantial. Rosenthal later summarised the routes as four channels, and described the whole business as covert communication, which is exactly the right phrase for it [15].

The first channel is climate. Warmth, proximity, eye contact, the microsecond of extra patience before moving on. This is the least deliberate of the four and the hardest to fake.

The second is input. Children a teacher expects more of get given harder material and more of it. This one is nearly invisible from inside the classroom, because from the child's point of view the work is just the work.

The third is output. How often you get called on, how long the teacher waits after asking you a question before giving up and moving to someone else, whether an incomplete answer gets a follow-up or a polite nod.

That third one is worth pausing on, because it is the channel a child can actually count. Wait time is a real, measurable quantity, and the difference between one second and four is the difference between being asked a question and being asked to fail at one.

The fourth is feedback. Not how much praise, but how differentiated it is. A child a teacher believes in gets told specifically what was wrong. A child a teacher has written off gets told "good try", which is the sound a teacher makes when they have stopped expecting an improvement.

Sarah Gentrup and colleagues traced that fourth channel directly in 2020, linking the feedback teachers actually gave to the achievement children actually reached [16]. John Darley and Russell Fazio had already described the general shape of the loop back in 1980, outside classrooms entirely: an expectation shapes a behaviour, the behaviour shapes the other person's response, and the response confirms the expectation that started it [17].

The loop is the thing to hold on to. It is not a teacher believing something and a child obliging. It is a circuit, and the return path is what makes it self-confirming.

Boris Eckstein and colleagues drew that circuit explicitly in 2026, surveying 85 elementary school teachers and 1,412 students to ask how far the label "behaviour problems" is earned by what a child actually does, and how far it comes from conditions that bias the teacher's perception [18]. Their model puts the teacher's experience of a child and the child's behaviour on opposite sides of a loop, each feeding the other, and sets the whole thing inside a frame of classroom conditions that nobody in the room chose.

A theoretical process model of pedagogical interactions in the classroom

Eckstein B, Grob U, Reusser K, Wettstein A. Teachers’ labeling of student behavior problems: a multiperspective study of teacher, student, and classroom conditions. Front Psychol.; 17:1704331. https://doi.org/10.3389/fpsyg.2026.1704331. FIGURE 1. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.

The loop, drawn by researchers who study it. The teacher's experience of a student and the student's behaviour sit on opposite sides of a circuit, each feeding the other, with the personal characteristics of both parties inside it and the conditions of the classroom around it. Nothing in this diagram is a single cause.

Teacher forms expectation

Climate: warmth and attention

Input: harder material

Output: more turns

Feedback: specific not vague

Student reads the signal

Student behaviour shifts

Miles Patterson's model of nonverbal communication describes why so much of this runs below deliberate control [19]. The teacher is not choosing to lean in. The teacher is leaning in.

The Children Can See It

Here is the part that tends to surprise people.

In 1987 Rhona Weinstein and Hermine Marshall asked children directly what they noticed about how their teacher treated different pupils [20]. The children could describe the differential treatment. Their awareness of it varied by age and by classroom, and in some classrooms it was very sharp indeed. The covert communication is only covert to the adults.

Think about what that means for your own school memories. If you had a sense of where you stood in a room, you were probably not imagining it, and you were probably not being told either.

Elisha Babad, Frank Bernieri and Robert Rosenthal pushed this further and got a result that still reads as slightly uncanny. In 1989 they showed people brief clips of teachers, without the context, and asked them to judge what those teachers expected of their students [21]. Strangers could do it. The title of the paper is "When less information is more informative", and the finding is that a very short exposure can be enough, because the giveaway is in the manner rather than the content. Two years later the same group showed that students themselves are reliable judges of a teacher's verbal and nonverbal behaviour [22].

So the picture is not one of a hidden influence working on unwitting children. It is closer to the opposite. The children generally know. Margaret Kuklinski and Rhona Weinstein later modelled how this plays out across development and found the pattern differs by classroom and by age [23], and recent longitudinal work has followed how stable a student's own reading of their teacher's expectations stays over time [24].

What a child does with that knowledge is the next question.

Grzegorz Szumski and Maciej Karwowski have argued that the route runs through academic self-concept, meaning the child's own sense of what they are capable of [25]. The expectation does not act on performance directly. It acts on what the child thinks they are, and performance follows from that.

That is worth sitting with, because it changes what you would do about it. If the route runs through self-concept, then the thing to protect is not a test score. It is what a child concludes about themselves from a year of small signals.

Related work traces the same path through self-efficacy and through the emotions a child brings to a subject [26]. Others have followed it into engagement, which is the plainest version of the idea: how much of themselves a student actually puts into the work [27].

A study of 583 students in urban and township junior high schools found achievement motivation sitting in the middle of the chain, between what students believed their teachers expected and what those students went on to achieve in Chinese, mathematics and English [28]. A separate line of work has tied perceived expectations to academic engagement in the same age group [29]. Not everyone reads it that way, though. One 2022 study took the deliberately provocative title "No more Pygmalion" and argued that a student's sense of mattering and their own self-efficacy carry more of the weight than the teacher's expectation does on its own [30]. That is a genuine disagreement about where in the chain the action happens, and it is not resolved.

That connects to a much larger literature on how external expectations become internal motivation, and it is the least mysterious part of the whole story. A child who is told, in a hundred small ways, that this is not for them will eventually agree.

The Trap: Accuracy Is Not Prophecy

Now the objection that reorganised the field, and it is the one thing almost every popular article on this topic gets wrong.

Teacher expectations correlate with student achievement. This is beyond doubt and easy to measure. Ask a teacher in September how a child will do, check in June, and the two will line up reasonably well.

Almost everyone treats that correlation as evidence of a self-fulfilling prophecy. It is not evidence of anything of the kind.

If you take one methodological point away from this article, make it that one. It is the difference between a claim about the world and a claim about a graph.

There is a much duller explanation available: teachers are reasonably good at judging who is doing well. A teacher predicting that a strong reader will still be a strong reader in June is not casting a spell. They are reading the situation correctly. The correlation would be exactly the same in a world where teacher beliefs had no causal power at all.

Lee Jussim has spent a career on this distinction, starting with a theoretical treatment in Psychological Review in 1986 [31] and then testing it directly. His 1989 paper pulled the three things apart: how much of the expectation-achievement link is self-fulfilling prophecy, how much is perceptual bias, and how much is plain accuracy [32]. He and Jacquelynne Eccles extended it in 1992 [33]. The answer, repeatedly, is that accuracy does most of the work.

He restated the general position much later, arguing that across social perception as a whole, accuracy dominates bias and self-fulfilling prophecy rather than the other way round [34]. You do not have to accept the strong form of that claim to accept the narrow one, which is unavoidable: a correlation between a belief and an outcome tells you nothing about which caused which. West and Anderson had said as much in 1976 and nobody had listened hard enough [9].

There is something slightly deflating about this, and the deflation is the point. A teacher looking at a child and thinking "this one will do well" is usually just a person being right about something. Most of the time nothing spooky is happening at all.

This is why the experimental studies matter so disproportionately. Only a study that plants a belief the teacher has no reason to hold can separate the prophecy from the accuracy. And those are exactly the studies whose pooled effect Raudenbush found to be small and dependent on timing.

There is a wider version of this trap, which is the tendency to be confident about a judgement whose accuracy you have never tested. Teachers are not unusual here. It is the same problem as the gap between feeling that you know something and actually knowing it, and it is related to why self-assessment of skill is so poorly calibrated in general.

Two Literatures Wearing One Name

Everything above is about experiments. There is a second body of work on teacher expectations that is much larger, much more recent, and reports much bigger numbers, and it answers a completely different question.

Instead of planting a false belief, this line of work finds teachers who naturally hold high expectations for everyone in their class and compares their students against the students of teachers who do not. In 2025 Christine Rubie-Davies and John Hattie pulled that literature together across seventy-six separate contrasts and reported a weighted average of 0.87 standard deviations [6].

They are large. Comparing teachers Elisha Babad classed as high versus low bias, the average across five contrasts is 0.92 standard deviations [6]. Comparing Rhona Weinstein's high versus low differentiating teachers, the average across twenty contrasts is 0.85. Comparing Rubie-Davies's own high versus low class-level expectation teachers, the average across fifty-one contrasts is 0.87, and reading progress alone reaches 0.97. The weighted average across all seventy-six contrasts is 0.87. For context, the same review cites Hattie's synthesis of nine meta-analyses at 0.58.

Set those against 0.078 and you can see why popular writing about this topic is such a mess. The two sets of numbers are ten times apart and they get quoted interchangeably.

The observational work is not thin, and it is worth being fair to it. Rubie-Davies and Hattie's earlier 2006 study took 540 students taught by 21 primary teachers and compared what those teachers expected in reading against what the students actually achieved, broken down across Maori, Pacific Island, Asian and New Zealand European groups [35]. That is a real measurement of a real gap between belief and performance. What it is not, and what no study of this design can be, is a demonstration that the belief created the gap.

Here is the thing that has to be said in the same breath as any of those large figures. Every one of them is an observational comparison between different teachers, not an experiment. Nobody randomly assigned teachers to hold high or low expectations. A teacher who expects a great deal of every child in the room is different from a colleague who does not in an enormous number of ways at once, and expectation is only one of them. They may plan differently, group differently, know their subject better, or have been given a different class in the first place. You cannot read those numbers as the causal effect of expectation any more than you can read the correlation between reading books and vocabulary as the causal effect of books.

You will find this confusion everywhere once you know to look for it. A page will describe the 1968 experiment, then reach for a number from the 2020s observational work to say how big the effect is, without noticing that the two sentences are about different studies of different things.

That is not a reason to dismiss the observational work. It describes something real about classrooms, and it is measuring something the experiments cannot reach, because no ethics committee is going to randomly assign a teacher to expect nothing of a class for a year. It is a reason to keep the two literatures in separate boxes.

What was comparedWhat it foundWhat it cannot show
The Oak School experimentAbout 0.35 standard deviations on measured abilityWhether the effect survives better measurement in the youngest grades
Nineteen induced-expectation experiments pooled0.078 standard deviations and not statistically significantThat the effect is absent. The studies differ far more than noise explains
The same nineteen with prior contact modelled0.134 at average contact and minus 0.157 per extra weekHow the slope behaves outside the range of contact the studies happened to cover
Thirty-five years of evidence reviewedReal but typically small and inclined to fadeWhich specific children in which specific rooms are affected
Seventy-six contrasts between high and low expectation teachers0.87 standard deviations on averageAnything causal. Nobody assigned teachers to their expectations

Where the Effect Stops Being Small

If the average is small, the average is also the wrong thing to look at. The most important sentence in this entire literature says so directly.

In 2005 Lee Jussim and Kent Harber weighed thirty-five years of research and reached four conclusions [36]. Self-fulfilling prophecies in the classroom do occur. They are typically small, they do not accumulate greatly across teachers or over time, and they may be more likely to dissipate than accumulate. The strongest self-fulfilling prophecies may occur selectively among students from stigmatised social groups. And teacher expectations may predict student outcomes more because those expectations are accurate than because they are self-fulfilling.

The second of those is where the harm lives, and it is the conclusion that gets left out of almost every popular summary. Average effects are reassuring. Concentrated effects are not.

The classic demonstration is Ray Rist's 1970 study, and it remains uncomfortable reading fifty-five years later [37]. Rist watched a single kindergarten classroom as an ethnographer. By the eighth day of school, before any academic work had been assessed, the teacher had permanently assigned the children to three tables. The assignment tracked social class closely: the children at the top table were cleaner, better dressed, more likely to have employed parents, more likely to speak standard English. Rist then followed those children forward and watched the table assignment turn into a fact about their schooling. Because this is one classroom observed by one researcher, it is a case study and not a measurement of how often this happens. It carries its weight by showing the mechanism in motion.

The statistical version arrived much later. Nicole Sorhagen followed children forward and found that early teacher expectations disproportionately affected the later high school performance of children from poor families [38]. Same shape as Rist. Much larger sample.

If you are reading this as a description of villains, stop. Nothing in this section requires a teacher to dislike a child, or to hold any view they would recognise as prejudice. That is the uncomfortable part.

Linda van den Bergh and colleagues went after the mechanism in 2010 with a study of 41 elementary teachers and 434 students [39]. They measured teachers' prejudiced attitudes two ways: by asking them, and with an implicit association test. What teachers said about their own attitudes related to nothing. The implicit measure related to their expectations and to the achievement gap in their classrooms. Whatever is going on here is mostly not happening at the level of opinions people would defend out loud.

That result has held up and been extended, and the extension produced the most specific number in this whole article.

Clark McKown and Rhona Weinstein worked with two independent datasets covering 1,872 elementary-aged children in 83 classrooms [40]. In ethnically diverse classrooms where the children themselves reported a lot of differential treatment, teachers' expectations of European American and Asian American students ran between 0.75 and 1.00 standard deviations higher than their expectations of African American and Latino students who had similar records of achievement [40].

Read that last clause again. Similar records of achievement. The gap sat in the expectation, not in the performance the expectation was supposedly based on.

The same paper carries the other half of the finding, and it is the hopeful half. In equally diverse classrooms where differential treatment was low, teachers held similar expectations for all students with similar records. The classroom, not the teacher's assumptions alone, decided whether the gap appeared at all.

Georg Lorenz has since studied the subtler forms this takes among teachers who hold no hostile views whatsoever [41].

The pattern is not restricted to one characteristic, and this is where the topic stops being a curiosity about a 1960s experiment.

Gender is the best documented. Joanne Rossi Becker was already recording differential treatment of girls and boys in mathematics classrooms in 1981 [42], and Jacquelynne Eccles and colleagues were working in the same decade on how classroom experience shapes what children come to believe they are good at [43]. Sarah Gentrup and Camilla Rjosk brought it up to date by examining how expectations feed into the gender gap in early achievement [44].

Aleksandra Gajda and colleagues went at the same question by watching. They logged 204 hours of observation across 34 classes of thirteen and fourteen year olds, then interviewed the 25 teachers who had taught those lessons [45]. The teachers knew gender stereotypes existed. Knowing did not keep the behaviour out of the observations.

Ask yourself what you would have to do to catch this in your own classroom. You would need a record of who you called on, how long you waited, and what you said when the answer was wrong. Almost nobody has that record.

Anneke Timmermans and colleagues have tested gender and minority background together as moderators, and pushed the outcome measure past achievement to what expectations do to a child's self-concept [46].

Disability and family income distort judgement the same way. Kim Smeets and colleagues studied 1,073 primary school children in grades four to six, drawn from 77 classes across 16 schools, alongside their teachers and their parents [47]. Both special educational needs and socioeconomic status pulled the adults' judgements of a child's cognitive ability away from what a cognitive ability test actually showed. Charlotte Schell and colleagues tested N = 76 German trainee teachers of average age 22.75 and found the same explicit-implicit split that van den Bergh found in serving teachers [48]. It arrives with them. It is not picked up on the job.

The list keeps growing, and the additions keep being things a child cannot change.

Richard Nennstiel and Sandra Gilgen took an intersectional approach and found body weight among the characteristics that shift a teacher's grading [49]. Even a child's accent moves what a teacher expects, in work going back to 1972 [50].

None of these characteristics has anything to do with what a child can do. All of them are visible in the first five minutes.

Put that next to the timing finding and you have the whole uncomfortable picture in one frame. The window in which an expectation does the most work is exactly the window in which a teacher has nothing to go on except what a child looks like, sounds like and arrives with.

The school matters too, not only the teacher. Orhan Agirdag and colleagues connected the effect to school composition, showing how segregation and expectation interact in mathematics achievement [51]. Birgit Heppt and colleagues used PISA data collected in Germany in 2018, with N = 2947 ninth-grade students, of whom 198 came from the most highly stigmatised minoritised groups and 445 from other minoritised groups [52]. Students in that first group reported a more discriminatory school climate than any other group did, and the study traced the path from that climate through teaching quality to how well those students adjusted.

Then, in 2020, an entire country ran the experiment by accident.

England, 2020

When the pandemic cancelled examinations in England, students could not sit the papers that would have decided their grades. Teachers were asked to supply centre assessment grades instead: their own judgement of what each student would have achieved. For one year, at national scale, teacher expectation was not a predictor of the outcome. It was the outcome.

Louis Magowan and colleagues treated this as a natural experiment and went looking for bias in teacher judgement across the demographic and socioeconomic characteristics of students, using national data from 2018 to 2020 as the comparison [53]. It is a rare thing in this field: a large, real, unplanned test of what happens when judgement replaces measurement.

It is worth being clear about why this is such a good test. Normally you cannot separate a teacher's judgement from the exam that checks it, because both exist. In 2020 only the judgement existed, and the model can estimate what the exam would have said.

Economists have been circling the same problem from another direction. Andrew Hill and Daniel Jones used quasi-experimental methods to get at self-fulfilling prophecies in classrooms without relying on either a planted label or a simple correlation [54]. And the pandemic disrupted expectation formation in a second way as well, by removing the classroom entirely for a period: Ariana Garrote and colleagues gathered 539 respondents and 83 teachers to ask what happened to teacher expectations, and to the parents suddenly standing in for the classroom, during emergency distance learning [55].

Similar accidental inductions turn up elsewhere once you look. The relative age effect is one. A child born in the last month of the school intake year is up to eleven months younger than a classmate in the same room, and the resulting difference in maturity gets read as a difference in ability. Geir Oterhals and colleagues examined the birth month, track choice and gender of all 28,231 students at upper secondary level in the Norwegian system, where the track decision is made at fifteen or sixteen, and found the relative age effect still shifting that choice a decade after the age difference stopped mattering [56]. Sofie Bolckmans and colleagues have traced the same pattern into achievement in sport [57].

Nobody labelled these children. The calendar did, and the adults around them read eleven months of extra development as ability.

If you were born in the summer and always felt slightly behind at school, that is not a memory to dismiss. It is one of the better documented findings in this whole area, and it has nothing to do with you.

One direction of this has been studied far more than the other, and it is not the direction that matters most.

The Mirror Image

Almost everything written about this topic is about high expectations. The literature is not.

The Pygmalion effect has two siblings, and the distinctions are worth getting straight because they are constantly muddled.

EffectWhose expectationDirectionWhat it predicts
PygmalionSomeone else's expectation of youHighRaised performance when the belief reaches you and changes how you are treated
GolemSomeone else's expectation of youLowDepressed performance through the same channels running the other way
GalateaYour own expectation of yourselfHigh or lowPerformance tracking self-belief without needing anyone else to hold it

The Golem effect is the one with the policy consequences, and it is barely covered anywhere outside the technical literature. It is also the reason the original experiment could never be run in reverse. You can tell a teacher that a randomly chosen child is about to bloom, and if you are wrong, the child has had a good year. You cannot tell a teacher that a randomly chosen child is about to stall. No ethics committee would approve it, and no researcher would propose it.

That asymmetry is not a small point about research ethics. It means the evidence base is systematically better on the half of the phenomenon that does less damage, and thinner on the half that does more. Any honest account of this topic has to say so.

This creates a real hole in the evidence. The direction of the effect that matters most for children is the direction nobody is allowed to test directly. What we have instead is the observational work above: Rist's tables, Sorhagen's follow-ups, the implicit attitude findings, the natural experiments. All of it points the same way, and none of it has the clean causal structure of the bloomer studies.

Does Any of It Last

There is a related question about whether any of this holds up when you follow the same children for years rather than months.

Faiza Jamil and colleagues have gone at it twice, tracking expectancy effects on mathematics achievement over time in one study [58] and on reading achievement in another [59]. Alena Friedrich and colleagues looked at what teacher expectancy does to mathematics achievement across a single school year [60].

Luis Soto-Ardila and colleagues went further down, to children solving individual arithmetic problems, in a study covering 1,420 students taught by 66 teachers across 48 schools in Spain [61]. They found a moderate correlation between expectation and performance, and one detail worth holding on to: the expectations ran higher than the results. Mathematics is a useful place to look, because it is the subject where a child's belief about whether they are "a maths person" hardens earliest and hardest.

Notice that the Jamil studies had to wait years for their answer. That is the practical reason this literature moves so slowly, and part of why the popular version got so far ahead of it.

Margaret Kuklinski and Rhona Weinstein built a path model on 376 children from the first through fifth grades and found that the strength of the effect depended on how visible the differential treatment was to the children themselves, with the direct effect on end-of-year achievement declining as children got older [23]. Older children see more and are moved less. That is the same shape as the prior-contact finding, running on a different clock.

There is one more asymmetry worth naming. The evidence that classroom prophecies fade rather than snowball comes from studies of ordinary variation. Alison Smith, Lee Jussim and Jacquelynne Eccles tracked whether self-fulfilling prophecies accumulate, dissipate or hold steady over time, and found dissipation rather than accumulation [62]. Stephanie Madon, Jussim and Eccles went looking specifically for the powerful version of the effect and reported that it lives in particular circumstances rather than in the general case [63]. Fading is what happens when a belief keeps meeting evidence. It is not obvious that a belief a child has internalised about themselves behaves the same way.

So which is it for you? If a teacher decided something about you at seven, does that belief still have hold of you now, or did it wash out somewhere in the twenty years since? The research cannot answer that for any one person, and you should be suspicious of anything that says it can.

Recent work outside education is a useful check on that. In 2026 a team led by Aryan Yazdanpanah ran a set of experiments on 111 participants who were exposed to painful heat, to videos of other people in pain, and to a demanding mental-rotation task, at three intensity levels each [64]. Before each one they saw a social cue that appeared to show how ten previous participants had rated it. The cue was randomised and had nothing to do with the actual intensity. Expectations and experience both shifted toward the cue, and the effect did not need any reward or punishment to sustain it. Same mechanism as the classroom. Opposite trajectory, because it persisted rather than fading. Whether a prophecy fades or entrenches appears to depend on what it is a prophecy about, and on whether the world keeps offering corrections.

Boot Camp

If you wanted to design a clean test of the timing finding, you would need an institution that hands instructors a group of complete strangers on a fixed schedule.

The Israel Defence Forces do this several times a year.

In 1982 Dov Eden and Abraham Shani published "Pygmalion goes to boot camp" in the Journal of Applied Psychology [65]. Instructors were given information about the command potential of incoming trainees. The information was randomly assigned. The instructors had no prior contact with any of these people, which is the condition Raudenbush's meta-analysis would identify as decisive two years later. The trainees labelled as high potential went on to perform better.

That study is often cited as proof that the Pygmalion effect is real and general. It is better read as proof of something narrower and more interesting: the effect shows up cleanly exactly where the theory says it should, in a setting where the expectation genuinely arrives before any first-hand knowledge does.

The workplace literature that grew out of this is substantial, and its headline number is startling. D. Brian McNatt gathered 17 studies covering 58 effect sizes and 2,874 people into a meta-analysis in 2000 and reported an average effect of 1.13 standard deviations [66]. That is roughly fourteen times the pooled figure from the classroom experiments.

Before you reach for that number as proof that expectations move mountains, read the moderators McNatt reports alongside it. The effect was stronger in the military than in civilian workplaces. It was stronger with men than with women. And it was stronger where low expectations had been held initially, which is to say where there was the most room to move.

The military point matters most, and it is the same point as Eden and Shani. Military training is the setting where instructors most reliably meet strangers on a schedule. It is the purest available version of the condition that the classroom meta-analysis identified as decisive, and it is where the effect is largest. Two literatures, forty years apart, pointing at the same variable.

David Trouilloud and colleagues have tested the mechanism in physical education, where a teacher's belief about a student's ability is unusually visible in how much that student actually gets to do [67]. Mehmet Altay and colleagues followed it into a real educational decision rather than a test score, looking at what drives students to pursue English-medium instruction [68].

1987: An Argument on Consecutive Pages

In the ninth issue of volume sixteen of Educational Researcher, published in 1987, Samuel Wineburg's article "The Self-Fulfillment of the Self-Fulfilling Prophecy" runs from page 28 to page 37 [69]. Robert Rosenthal's reply, "Pygmalion Effects: Existence, Magnitude, and Social Importance", starts on page 37 and runs to page 41 [70].

You can turn one page and go from the case for the prosecution to the case for the defence. There is no better single object for understanding why this argument lasted so long.

Wineburg's charge, roughly, is that the field's belief in the Pygmalion effect was itself a self-fulfilling prophecy: people wanted it to be true, the original evidence was weaker than its reputation, and the reputation had done the rest. Rosenthal's reply is not that the effect is enormous. It is that the effect exists, that its magnitude is being systematically understated by critics who prefer significance tests to effect sizes, and that even a modest effect matters when it operates on every child in every classroom.

Both men are partly right, which is the least satisfying and most accurate thing to say about it. If you want a single object that explains why this argument ran for fifty years, it is those consecutive pages. Twelve years later Herman Spitz wrote the history of the controversy itself, from the outside, and the fact that the dispute had by then generated its own historiography tells you something about how long it ran [71].

Rosenthal had made his broader case a few years earlier as well, arguing that interpersonal expectancy effects seen across three decades of work formed a coherent phenomenon [72]. Claude Goldenberg came at it from the opposite side, arguing from close case knowledge that expectations run into limits imposed by everything else happening in a classroom [73].

Can Expectations Be Changed on Purpose

If low expectations do harm, the obvious question is whether you can raise them deliberately.

Hester de Boer, Roel Bosker and Margaretha van der Werf reviewed the intervention studies and asked exactly that [74]. The paper sits behind a publisher block that refuses automated access, so this article is not going to put a number on their summary effect. What can be said is that a body of intervention work exists, that it is aimed at raising teacher expectations and student achievement together, and that the authors considered the evidence worth synthesising.

Francesca López has written about altering the trajectory of the self-fulfilling prophecy in teacher preparation [75]. Rianne Bosman and colleagues have tested a structured reflection approach aimed at improving teacher-child relationships [76]. Gintautas Katulis and colleagues took the protective side, following 540 Lithuanian students in grades five to seven and surveying them twice about 5 months apart, and found that a positive classroom climate buffered temperamentally vulnerable children against rising loneliness [77].

The honest status of this work is promising and under-tested. Nobody should read it as a solved problem, and this article is not going to tell you that a training day fixes anything.

What is worth noticing is where the theory says an intervention would do the most good. If the effect concentrates in the window before a teacher has their own evidence, then the highest-value target is not the classroom in March. It is the handover in August: what gets written in a transfer file, what a streaming decision is based on, what a teacher is told about a child they have not met. That is where a belief is doing the most work and facing the least resistance.

The Version That Arrives as a Spreadsheet

The Oak School study worked by handing teachers a plausible-looking prediction about a child's future, generated by a process the teachers did not understand and could not check.

That description fits something being built right now.

Polygenic scores, which compress many small genetic associations into a single number, are already sold direct to consumers and are being discussed as inputs to education under the heading of precision education. Lucas Matthews and colleagues named the problem in 2024 and called it the polygenic Pygmalion effect [78].

They did not stop at naming it. They ran two studies, each randomising 1,188 students to one of four conditions: a low-percentile polygenic score for educational attainment, the same score with mitigating information attached, the same score with exacerbating information attached, or a control [78]. That design treats the score exactly as Rosenthal and Jacobson treated their bloomer list, as an induction whose effects can be measured. Asking whether the damage can be mitigated is the right question to be asking before the thing is deployed rather than after.

The structural resemblance to Oak School is close enough to be uncomfortable. A number arrives from an authoritative-sounding source. It makes a claim about a child's future rather than their present, so nothing the child does this week can contradict it. The teacher has no way to evaluate it. And crucially, it lands before the teacher has formed their own view, which is precisely the condition under which the experimental literature says an expectation moves a child.

Mayli Mertens and Angus Clarke have written about genetic tests functioning as prophecy more generally, and about the difference between prophecies that defeat themselves and prophecies that fulfil themselves [79]. Rémy Furrer and colleagues have examined how prenatal testing anchors expectations about a child before birth [80].

You do not need to believe polygenic scores predict anything much to worry about this. The Oak School bloomer list did not predict anything either. That was the whole design.

None of this requires anyone to behave badly. Rosenthal and Jacobson's teachers were not villains. They were given information they had no reason to doubt, at the one moment when they had nothing to weigh it against.

Time to put it all in one place.

What Is Actually True

After sixty years, here is what the evidence supports.

Teachers do treat children they expect more of differently, and the differences run through warmth, difficulty, opportunity to speak and the quality of feedback. This is not contested. It has been measured directly, it is visible to outside observers from very short samples of behaviour, and the children themselves can describe it.

Experimentally induced expectations produce, on average, a small effect on measured ability, and the effect shrinks the longer the teacher has already known the child. Pooled with no other information the average is not statistically distinguishable from zero. Account for prior contact and the picture sharpens into something real and steeply time-limited.

Self-fulfilling prophecies in classrooms are real. They are typically small, they do not accumulate much across teachers or years, and they are more likely to fade than to build.

The larger effects concentrate among children from stigmatised groups, and that is where the practical harm lives. This is the part of the literature with the strongest evidence and the least popular coverage.

Much of the correlation between what a teacher expects and what a child achieves is accuracy rather than causation, and any argument that ignores this is not an argument about the Pygmalion effect at all.

The effect is not confined to schools. It shows up in military training, in workplaces, in sport, and in domains as far from the classroom as the judgement of pain.

And the timing finding is worth restating as plainly as possible, because it is the one that changes what anybody would actually do. The most valuable moment to intervene is not in the classroom in March. It is in the file in August.

And four things remain genuinely open. Whether the Oak School result would survive better measurement in the youngest grades. Whether the effect is large enough to be a policy target. Whether an expectation a child has internalised behaves like the ones that fade or like the ones that entrench. And whether expectations can be changed on purpose in a way that lasts.

The self-help version of this topic tells you to believe in people and watch them flourish. The research says something less convenient and more useful. Belief is not a force that flows into a person. It is a set of small behavioural differences that reach someone through a channel, and the channel is widest at the beginning, when nobody has any evidence yet.

Which is why the sentence that matters is not the one about what a teacher believes about you.

It is the one about what a teacher believes before meeting you.

Frequently Asked Questions

What is the Pygmalion effect?

The Pygmalion effect is the finding that a teacher's expectations about a student can influence how that student performs, by changing how the teacher treats them. It is named after the Greek sculptor who fell in love with his own statue. It was introduced by Robert Rosenthal and Lenore Jacobson in the 1960s, and the evidence for it turns out to be real but considerably smaller and more conditional than the popular version suggests.

Is the Pygmalion effect real, or was the original study debunked?

Both halves of that question are wrong. The original study was heavily criticised, particularly on how it measured ability in the youngest children, and many attempts to replicate it failed. But the phenomenon itself survived. A 2005 review of thirty-five years of evidence concluded that classroom self-fulfilling prophecies do occur, that they are typically small, and that they are more likely to fade than to accumulate. "Debunked" is as inaccurate as "one of psychology's most powerful effects".

How big is the teacher expectation effect, actually?

It depends entirely on what is being compared, which is why the numbers you see vary so wildly. Pooling nineteen experiments that planted a false expectation gives an average of about 0.078 standard deviations, which is not statistically distinguishable from zero. Account for how long the teacher had already known the children and the effect at average prior contact is 0.134, falling by about 0.157 for each additional week of prior contact. Studies comparing high-expectation teachers with low-expectation teachers report much larger figures, averaging around 0.87, but those are observational comparisons between different teachers rather than experiments and cannot be read causally.

What is the difference between the Pygmalion, Golem and Galatea effects?

The Pygmalion effect is someone else's high expectation raising your performance. The Golem effect is the same machinery running in reverse, where someone else's low expectation depresses it. The Galatea effect is your own expectation of yourself affecting what you achieve, with no one else required. The Golem effect matters most in practice and is studied least, partly because deliberately giving a teacher low expectations of a real child could never be approved as an experiment.

Can a teacher's low expectations actually hold a student back?

The evidence points that way, though it is observational rather than experimental for the ethical reason above. Studies following children forward have found that early teacher expectations disproportionately affect the later school performance of children from poor families, and that teachers' implicit attitudes, rather than the views they would state out loud, relate both to their expectations and to achievement gaps within their classrooms.