Introduction
You hand something in. A test, a report, a piece of code. A few days later it comes back with a number on it. Maybe a 6 out of 10. Maybe a ranking against the rest of the team. Maybe a line that says "great work, you're a natural".
Almost everyone agrees that this is how people get better. You try, somebody tells you how it went, you adjust. Teachers are trained in it, managers are sent on courses about it, and apps are built around streaks and scores.
It works on average. That part is not in doubt.
But in 1996 two psychologists, Avraham Kluger and Angelo DeNisi, pulled together the experimental literature on feedback and found something that should have been front-page news in every school and HR department.
In more than a third of the comparisons they collected, the people who received feedback performed worse than comparable people who received none at all [1].
Not a little worse by accident. Worse in a way that sampling error could not explain, and that could not be blamed on the feedback being bad news.
That is the feedback intervention effect. It is not the claim that feedback fails. It is the claim that feedback is a gamble whose odds depend on what the message does to your attention, and that we have been averaging the losses away for a century.
The Century Nobody Counted the Losses
Feedback has been at the centre of learning theory for as long as there has been learning theory.
Edward Thorndike's law of effect, restated in a 1927 paper, held that responses followed by satisfying consequences are strengthened and those followed by annoying ones are weakened [2].
Knowing whether you got it right is the simplest satisfying or annoying consequence there is.
By mid-century the idea had hardened.
A 1956 survey of the research on knowledge of performance set out a list of tentative generalisations, and the overall tone was that knowing your results helps you learn and keeps you trying [3].
There were warning signs. Studies showing feedback making things worse had been published since early in the century. They were treated as noise.
A 1979 review of feedback in organisations came closer to the problem. Daniel Ilgen and colleagues argued that feedback does not act on people directly.
It has to be perceived, then accepted, then acted on, and it can fail at any of those three steps [4].
That turned feedback from a switch into a message, and messages can be misread.
The two biggest ideas on that timeline, that feedback can backfire and that the reason is attention, arrived late. Several findings built on them have since been questioned.
1996: The Number That Should Have Changed Everything
Kluger and DeNisi did something that sounds obvious and had not been done properly. They collected every study they could find that compared people who received feedback with people who did not, on a measure of actual performance, and pooled them.
The result was 607 effect sizes drawn from 23,663 observations [1].
On average, feedback improved performance by about 0.41 of a standard deviation. That is a respectable effect. In education it would count as a good intervention.
Then they looked at the spread.
More than a third of the effects were negative [1].
In those studies, the feedback group ended up behind.
This was not a handful of odd results in the tail. It was not chance, it was not about whether the feedback was positive or negative, and none of the theories of the day predicted it.
Writing about the finding two years later, they put it plainly: feedback improved performance on average but reduced it in more than one third of cases, which ran against the common belief that it usually helps [5].
Think about what that means for you. If you give students a score on every assignment, or run a monthly performance chat, the research does not say you are wasting your time. It says that roughly one time in three, the same kind of intervention made people worse, and the average hides which one you are running.
They also proposed an explanation, and called it preliminary in the title of the paper. That word matters.
It Was Never About Good News or Bad News
The obvious guess is that the harmful third was mostly criticism. Tell people they did badly, they get discouraged, performance drops.
That guess is wrong, or at least far too simple.
In the 1996 analysis the sign of the feedback did not explain which interventions backfired [1].
Positive feedback could hurt. Negative feedback could help.
Later work kept finding the same slipperiness.
A 2018 meta-analysis of 78 studies on negative feedback and intrinsic motivation found that negative feedback had no effect on motivation when compared with neutral feedback or none at all [6].
Negative feedback only looked damaging when compared with praise. The same review found negative feedback was less demotivating when it included instructions on how to improve, when it was judged against a clear standard rather than against other people, and when it was given in person.
A study by Dina Van Dijk and Kluger added a twist.
Across samples of 315 and 318 people, positive feedback raised self-reported motivation relative to negative feedback on tasks that called for creativity (d = 0.43), and lowered it on tasks that called for vigilance and attention to detail (d = -0.33) [7].
Performance moved the same way, in smaller samples of 55. On a proofreading-style task, being told you are doing well can make you relax at exactly the wrong moment.
So the valence of the message is not the lever. The lever is something else, and Kluger and DeNisi thought they knew what it was.
Where Your Attention Goes When Someone Hands You a Number
Feedback intervention theory starts from an older idea in psychology called control theory. The basic unit is a loop.
You have a standard, you compare where you are against it, and you act to close the gap [8].
A thermostat does this. So, the theory goes, do you.
Kluger and DeNisi's addition was that these loops are stacked. At the bottom are loops about the details of the task: which step went wrong, which answer was off. Above those are loops about effort on the task: am I trying hard enough, should I speed up.
At the top are loops about you: what does this say about my ability, how do I compare, what will people think [1].
Attention is limited. When feedback arrives, it pulls attention toward one of these levels.
The central claim of the theory is that feedback gets less effective as attention moves up the stack, closer to the self and further from the task [1].
This explains why sign alone does not predict harm. "You got question four wrong because you dropped a minus sign" is negative, but it keeps you at the bottom of the stack. "Brilliant, you're clearly one of the strongest in the class" is positive, but it sends you straight to the top. Now you are thinking about staying there, not about the next problem.
It also fits a much larger literature on goals.
Decades of goal-setting research show that feedback works best when it is paired with a specific goal, because the goal tells you what to do with the information [9].
Feedback with no goal attached leaves you to supply one yourself, and the goal people supply by default is often about how they look. If you want the background on that side of the story, our article on why specific and hard goals beat "do your best" covers it.
So feedback that points at the self should do worse than feedback that points at the work. The cleanest classroom test of that prediction was run a decade before the theory existed, in Israeli primary schools.
The Classroom Experiment: A Grade, a Comment, or Both
Ruth Butler was interested in a simple question. When you give a child feedback on a piece of work, does the form of that feedback change how interested they are in doing more?
In a 1986 study with Mordecai Nisan, Butler gave children either no feedback, task-related comments, or grades after they completed tasks. Comments produced the best later performance and the most sustained interest.
Grades did not [10].
A 1987 study pushed further.
In it, 200 students in fifth and sixth grade worked on creative thinking tasks across three sessions, and after the first two they got individual comments, numerical grades, standardised praise, or nothing [11].
By the third session, interest and performance were highest in the comments group, at both levels of achievement. Children who got grades or praise were more likely to explain their results in terms of ability and how they compared with others. Praise and grades, two things that look opposite, behaved alike. Both pointed at the self.
Then came the study everyone in education research eventually cites.
In 1988 Butler randomly assigned twelve classes to receive grades, comments, or both, and tracked 132 children of high and low achievement across the sessions [12].
The comments group kept its interest and improved. The grades group did not. And the group that got both a grade and a comment looked like the grades group. Adding the grade wiped out the benefit of the comment.
If you have ever handed back marked work, you know why. The eye goes to the number, then to the number on the next desk. The comment is read, if at all, as an excuse for the number.
Why the Grade Swallows the Comment
Later studies found the same shape with older students and more careful designs.
In a 2009 experiment with university students, detailed written feedback on an essay was most effective when it came on its own, without a grade and without praise attached [13].
Whether students thought the comments came from a computer or from their instructor made little difference. What mattered was whether a score came with them.
A 2021 study of 464 students at university looked at what happens inside that moment [14].
The students wrote an essay, received feedback, and then revised. Receiving a numeric score directly predicted worse performance on the later essay exam and more negative emotion, and negative emotion in turn carried part of the effect.
A related 2016 study traced the drop in children's intrinsic motivation after negative feedback to their ability self-concept, in other words to what the feedback told them about how good they were rather than about the work [15].
The biggest summary of this literature is a 2019 set of meta-analyses on grades and comments in primary and secondary school [16].
Compared with no feedback at all, grades raised achievement but lowered motivation. Compared with comments, grades did worse on both. So a grade is better than silence for getting the next test score up, and worse than a good comment for almost everything.
But Grades Are Not Poison Either
It would be easy to stop there and conclude that grades are harmful. The more recent evidence does not support anything that clean, and that matters.
A 2025 study followed 99 students in secondary school across four classroom groups as grades were shown in some phases and hidden in others, while written comments stayed the same [17].
When grades first appeared, performance dipped and negative emotions rose, just as Butler would predict. But the effect faded over time. The students adapted.
A 2021 study used a regression discontinuity design to ask what a low midterm grade does to college students in the rest of the course [18].
Students just below a grade boundary improved more afterwards than students just above it. They did not do it by cramming for the final. They did it by picking up points on low-stakes work, such as participation, reading quizzes and in-class exercises.
And a large 2024 randomised trial with 2,736 students in ninth grade across 29 schools tested a grading system built around standards, formative feedback and the chance to be reassessed for full credit [19].
The new system raised end-of-course algebra and geometry scores by 0.33 standard deviations. The trial tested a bundle, so no single part can take the credit. But one thing that changed was what a grade meant: a marker on the way to mastery that you could revise, rather than a verdict.
That fits the theory. A grade that says "not there yet, here is what to do" keeps attention on the task. A grade that says "this is where you rank" sends it to the self. Our piece on what competition and social comparison do to learning covers the ranking side.
"You're So Smart": Praise Aimed at the Person
The praise studies are where the feedback intervention effect reached a mass audience, usually without its name attached.
In 1998 Claudia Mueller and Carol Dweck reported six studies with children [20].
After an easy set of problems, some children were praised for their intelligence and others for their effort. Then everyone got a much harder set on which they did poorly. The intelligence-praised children enjoyed the tasks less, persisted less and performed worse afterwards than the effort-praised children. They were also more likely to describe intelligence as a fixed trait.
A 1999 follow-up with young children found that person-focused criticism and person-focused praise both produced more "helpless" responses after a setback than praise aimed at the process, such as the strategy the child had used [21].
Praise that sounded kind and criticism that sounded harsh did the same damage when both were about who the child was.
Both results sit neatly inside Kluger and DeNisi's framework. "You're so smart" is feedback aimed at the top of the stack. It turns the next task into a test of that label.
The idea underneath came from a 1988 model in which a person's belief about whether ability is fixed steers them toward one of two goals: proving how able they are, or getting better [22]. A 2006 brain-recording study showed what that difference looks like at the moment feedback arrives [23].
Students who believed intelligence was fixed reacted more strongly to being told they were wrong. Then they showed less sustained memory-related activity while the correct answer was on the screen, and they corrected fewer of their errors on a surprise retest. Their attention went to the verdict. It did not stay for the lesson.
A 2002 review of praise research reached a more cautious conclusion.
Praise can undermine intrinsic motivation, enhance it, or have no effect at all, depending on whether it seems sincere, what it credits the success to, and whether it invites comparison with others [24].
That is worth keeping in mind before you read any headline about praise.
Eddie Brummelman and colleagues found that adults tend to give children with low self-esteem exactly the kind of praise most likely to backfire: person praise and inflated praise, the "that's not just beautiful, that's incredibly beautiful" kind [25]. Inflated praise then led those children to avoid the challenges they most needed, a loop a 2016 review described as self-sustaining [26].
The same logic shows up outside the lab.
In a 2021 classroom study, 63 students learning English at a junior high school in Iran received effort praise, intelligence praise or neither across 14 speaking sessions, and effort praise strengthened their growth mindset, their sense of competence and their willingness to speak up [27].
The Praise Studies, Taken Apart
This is where the story gets less comfortable, and where most popular retellings stop.
In 2019 Yue Li and Timothy Bates ran a close replication of the 1998 praise design with 624 children in China aged 9 to 13 [28].
The first study copied the original procedure. The praise manipulation was linked to performance on a moderately difficult test after failure, but only barely (p = .049), and it had no reliable effect on any of the eight motivation and attribution measures the original paper used. The average p value across those measures was .48. Two further studies added an active control condition and harder material. Neither found an effect of intelligence versus effort praise.
That is not a small wobble. The motivational outcomes were the heart of the original claim.
The broader mindset literature has had a similar reckoning.
A 2023 set of meta-analyses covered 63 studies with 97,672 students and found major weaknesses in design and reporting, signs of publication bias, and an overall effect of growth mindset interventions on achievement of d = 0.05, which was not significant after correcting for that bias [29].
Studies run by authors with a financial stake in the interventions reported larger effects.
The general principle, that feedback aimed at the self tends to do worse than feedback aimed at the task, rests on a far larger base than the praise experiments: grades, rankings, workplaces and the 1996 meta-analysis itself. The dramatic version, that one sentence of intelligence praise measurably damages a child, is contested, with Dweck and colleagues on one side and Li and Bates and the 2023 meta-analysts on the other. For how proving yourself differs from improving, see our article on mastery goals versus performance goals.
When Watching Yourself Breaks the Skill
There is a second route by which feedback can hurt, and it has nothing to do with self-esteem. It has to do with attention to your own movements.
Research on choking under pressure makes the mechanism concrete.
Roy Baumeister proposed in 1984 that pressure increases conscious attention to the process of performing, and that this attention disrupts skills that are overlearned and automatic [30]. Later experiments by Sian Beilock and Thomas Carr supported this explicit monitoring account of choking for skilled performers [31].
Scores do exactly this. If you are a decent golfer and someone starts reading out your putting accuracy after every stroke, you may start thinking about your wrists. Thinking about your wrists is bad for putting.
The OPTIMAL theory of motor learning holds that performance and learning improve when attention is directed outward to the effect of a movement rather than inward to the body, and when learners feel some autonomy and expect to succeed [32].
Ranks, Neighbours and the Boomerang
Comparing people with each other is one of the most common forms of feedback, and one of the most unpredictable.
A 2007 field experiment on household energy use shows how the same number can pull in two directions [33].
Households were told how their electricity use compared with their neighbourhood's average. High users cut back. Low users, learning that they were below average, used more. The researchers called this the boomerang effect. Adding a simple sign of approval, a smiling face for households below average, removed the boomerang. The comparison alone gave people a standard to drift toward. The smiley told them which way was good.
In education, relative feedback has produced gains too.
When one Spanish high school told students for a single year whether they were above or below the class average, and by how much, grades rose by about 5 percent [34].
Who is hearing the comparison matters.
A 2023 study gave 152 students false feedback telling them how they had performed compared with others on an attention task [35].
Students who were chronic procrastinators showed weaker brain signals linked to attentional control and error monitoring than other students after negative comparisons, but the difference disappeared when the comparison was positive. The authors suggested that the same ranking may motivate some people and knock others off the task.
This is not a paradox. A rank says a little about the task and a lot about the self, and whether it helps depends on which part you hear. Our article on stereotype threat shows the same threat to the self costing performance in a very different setting.
The Kind of Feedback That Almost Always Works
After all that, it would be natural to feel wary of feedback. That would be the wrong lesson. There is one kind that is about as reliable as anything in psychology.
It is the correction of a specific error, given after you have tried.
In a 2005 series of experiments on learning words, telling people the correct answer after they got an item wrong increased their final retention of those items by 494 percent [36].
Feedback after correct answers made much less difference. Whether the correction came immediately or after a delay barely mattered.
Memory researchers have also found that errors made with high confidence are especially likely to be corrected when feedback arrives, a result known as the hypercorrection effect [37]. A 2017 review of learning from errors concluded that making mistakes and then receiving corrective feedback is good for learning, and particularly for mistakes a person strongly believed were right [38].
This kind of feedback lives at the very bottom of Kluger and DeNisi's stack. It says nothing about you. It says: this answer, not that one.
It has a catch, though.
A 1991 meta-analysis of 58 effects from 40 reports on feedback in test-like situations found that the benefit depended heavily on whether learners could see the answer before they had committed to their own [39]. If you can peek, feedback becomes a shortcut. This is the same reason self-testing works so much better than rereading, which our article on the testing effect explains in detail.
Even corrective feedback is not universal.
A 2016 experiment with 108 children learning a type of maths problem found that feedback helped children who did not yet know a correct strategy, but hurt children who had already been taught one [40].
For the second group, a stream of right-and-wrong signals seems to have pushed them to abandon a good method and try things at random. Sometimes struggling first and hearing the answer later is the better route, which is the argument made in our article on productive failure.
More Feedback, Less Learning?
One of the strangest claims in this field is that giving people less feedback can help them learn more.
The claim comes from motor learning.
In a 1984 review, Alan Salmoni, Richard Schmidt and Charles Walter argued that knowledge of results works partly as guidance [41].
When you get it after every attempt, performance during practice improves, because you keep correcting. But you come to lean on it. Take it away, and the skill turns out to be weaker than it looked. This became known as the guidance hypothesis.
A 1990 paper reported three experiments in which reducing how often learners received knowledge of results improved their retention of a motor skill, even though it made practice look worse [42]. Schmidt and Robert Bjork later argued that this was one example of a wider pattern across motor and verbal learning. Conditions that make practice look good often produce weak long-term learning, and conditions that slow practice down can produce stronger learning [43].
That is the core of what is now called desirable difficulties.
A 2021 study of a soccer throw found an inverted U: learners who received feedback on 50 percent of their trials, or who chose when to ask for it, retained the skill better than those who got it on every trial or on only 25 percent [44].
It is a neat story, and it is in the textbooks. It is also now in serious doubt.
A 2022 meta-analysis gathered the experiments on reduced feedback frequency and found no reliable effect on either learning or performance [45].
The authors described the studies as severely underpowered and the results as highly variable, with the source of that variation still unexplained. They concluded that the guidance hypothesis, as usually stated, is not supported by the evidence as it stands.
So the textbook and the meta-analysis disagree. Until larger, preregistered studies settle it, the fair position is that feedback frequency matters less clearly than was once taught.
Does Timing Matter?
The same thing has happened with timing.
A 1988 meta-analysis of 53 studies on feedback timing found a split that has confused teachers ever since [46].
Applied studies using real classroom quizzes usually found immediate feedback more effective. Experimental studies of learning test content usually found delayed feedback more effective. The authors traced much of the difference to how the studies were designed.
Then came a genuinely new kind of study.
In ManyClasses 1, researchers ran the same experiment in 38 college classes across different subjects, institutions and teachers, comparing feedback given immediately with feedback delayed by a few days [47].
The overall effect was 0.002, with a 95 percent interval from -0.05 to 0.05. None of 40 preregistered moderators showed a credible effect either.
That is about as close to zero as a result gets. For everyday coursework, the timing of feedback seems to matter much less than teachers are often told.
At Work: Appraisals, Dashboards and Algorithms
Kluger and DeNisi were organisational psychologists, so the workplace is where their result should have landed hardest.
A 2005 meta-analysis of 24 longitudinal studies of multisource or "360-degree" feedback found that improvements in ratings from direct reports, peers and supervisors over time were generally small [48].
Improvement was most likely when the feedback clearly showed a need for change, and when the person reacted positively, believed change was possible, and set goals.
A 2020 study shows why those conditions are hard to meet.
In a survey of 382 participants, all of them managers, those who had received negative feedback judged it as less accurate than positive feedback and judged the person who gave it as less qualified to do so [49].
In role-played reviews with 117 pairs of managers, disagreement about past performance was larger after the conversation than before it. One manager, defending this pattern in a group debrief, said: "We are the best there is. If we get negative feedback for something bad that happened, it probably wasn't our fault!"
That sentence is feedback intervention theory in one breath. The conversation was meant to be about the work. It became about the self, and the self defended itself.
The same study found the one variable that most predicted whether people accepted feedback and intended to change: how much the conversation focused on future actions rather than on the past [49]. A 2025 diary study of 212 people in full-time jobs, across 779 daily surveys, found a matching pattern from the other direction [50].
Task-focused feedback produced better reactions than self-focused feedback, because it was felt as less of a threat to self-worth.
Healthcare shows how variable feedback is at scale, because clinical behaviour is logged anyway. A Cochrane review of audit and feedback in healthcare, updated in 2012, found that it generally produced small improvements, and that feedback worked better when baseline performance was low, when it came from a supervisor or colleague, and when it was given more than once [51]. A 2021 evaluation of comparative feedback to 316 general practices in England, against 130 control practices, found that opioid prescribing fell in the feedback practices while it kept rising elsewhere, a difference corresponding to about 15,000 fewer patients prescribed opioids [52]. The effect weakened once the feedback stopped.
That is a real result. It is also one programme in one region.
A 2023 review of feedback to ambulance staff found a pooled effect on care quality of d = 0.50, but with an I² of 99 percent, meaning the studies disagreed hugely about how much it helps [53].
Algorithms change who seems to be judging you.
In a 2021 field experiment, AI-generated performance feedback improved employees' results because it was more accurate and consistent, but telling employees that the feedback came from an AI reduced their productivity, an effect that was weaker for longer-serving staff [54]. A 2024 study found the reverse for one group: employees who feared losing face responded better to negative feedback from an AI system than from a person, because it triggered less rumination and more motivation to learn [55].
The source of the message changed what it meant to the self.
Even the person giving feedback gets pulled toward the self.
A 2022 experience sampling study found that leaders high in empathy felt less attentive and more distressed after giving negative feedback to their staff [56]. It is hard to keep a conversation on the task while you are managing your own discomfort.
How Big Is Feedback, Really?
You have probably heard that feedback is one of the most powerful influences on learning.
That phrase comes from John Hattie and Helen Timperley's 2007 review, which added in the same breath that the impact can be positive or negative [57].
Their model sorted feedback into four levels: about the task, about the processes behind it, about self-regulation, and about the self. The fourth level, they argued, is the least effective, which is Kluger and DeNisi's result in a school uniform.
The size claims around this are often inflated.
A 2020 meta-analysis designed to update those estimates covered 435 studies, 994 effect sizes and more than 61,000 people [58].
That analysis found a medium average effect of d = 0.48. But it also found so much variation that feedback could not be treated as one consistent treatment. The information content of the message made a large difference, and feedback helped cognitive and motor skills more than motivation and behaviour.
A 2023 meta-analysis of automated writing feedback, covering 20 studies and 2,828 students, found a similar picture: a medium average effect of g = 0.55 with substantial heterogeneity [59].
There is a live methodological dispute here too.
Critics argue that effect sizes from very different studies cannot be ranked into league tables, because they depend on choices such as the test used [60]. Others argue that an effect Cohen's rules call small is often large for a real school intervention [61].
Either way, no single feedback number is a fixed property of feedback.
Read down the middle column and the pattern the 1996 paper proposed is still visible, even with its contested rows included. The rows that point at the task tend to help. The rows that point at the self tend to disappoint.
The formative assessment movement in schools grew out of this.
A 1998 review by Paul Black and Dylan Wiliam argued that strengthening the frequent feedback students get about their learning produces substantial gains [62], and a 2008 review of formative feedback concluded that it should be non-evaluative, supportive, timely and specific [63].
Non-evaluative is the key word. It is another way of saying keep it off the self.
What the Evidence Says About Giving Feedback
None of this is a checklist. It is a way of looking at any feedback you give or receive and asking where it sends attention.
Some of the most striking evidence concerns who is receiving the feedback and what they fear it means.
In a 1999 pair of studies, Black students who received critical feedback on an essay on its own rated the reviewer as more biased and were less motivated than White students [64].
When the same criticism came with a statement that the reviewer used high standards and believed the student could reach them, Black students responded as positively as White students. The criticism did not change. What it said about the student did.
A 2018 meta-analysis on negative feedback reached a compatible conclusion from a different direction: telling people how to improve took much of the sting out of negative feedback [6]. Kluger has argued for going further and replacing some feedback with feedforward, conversations about what has worked for a person in the past and how to repeat it [65].
For students, recent research has moved from what teachers say to what students do with it.
A 2021 review of feedback interventions on written work screened 19,065 papers and analysed 58, and concluded that feedback works when students are motivated and able to use it, and that self-determination theory fits the evidence well [66]. See our article on self-determination theory for why feeling controlled drains motivation. A 2021 survey of 308 students in an English studies course found that, at the less selective of two universities, what students actually did with feedback predicted their exam results, and teacher feedback mattered mainly through that behaviour [67].
Two recent reviews of 14 prominent feedback models proposed an integrated framework with five parts: the message, its implementation, the student, the context and the people involved [68] [69]. The words you choose are just one part of that picture.
What Is Settled and What Is Not
Some of this is firm. Feedback helps on average, by roughly 0.4 to 0.5 of a standard deviation in the large meta-analyses of 1996 and 2020, and in the 1996 analysis more than a third of effects were negative. The sign of the message does not decide which. Correcting a specific error after someone has tried is among the most reliable tools in learning science, provided the answer was not visible first.
Feedback aimed at the self, whether a grade, a rank or a label, tends to do worse than feedback aimed at the work. That rests on the most evidence, but its edges are softer than the classic studies suggest, because grade penalties can fade and grading built around revision can raise achievement.
Three claims remain open: that one sentence of intelligence praise measurably harms a child, that less frequent feedback improves motor learning, and that feedback can be ranked against other school interventions by effect size. The timing debate, meanwhile, looks overstated for ordinary coursework.
The theory itself is still, in its authors' word, preliminary. Its core idea about attention fits the pattern described here, from grades to rankings to workplace reviews, but the full theory has never been tested as a whole.
Conclusion
The feedback intervention effect is not an argument against telling people how they are doing. It is an argument against assuming that telling them is automatically good.
For most of a century the studies where feedback made people worse were set aside as flukes. Counted properly, they were more than a third of the evidence. The explanation that has lasted is about attention. A message that points you at the work gives you something to do. A message that points you at yourself gives you something to protect.
That is why a grade can cancel a comment. It is why a ranking can push one person up and another person off the task. And it is why the most reliable feedback of all is also the plainest: this one was wrong, here is the right one, try again.
The next time a number comes back on something you made, notice where your mind goes first. If it goes to what the number says about you, you have just watched the effect happen.
Frequently Asked Questions
What is the feedback intervention effect?
The feedback intervention effect refers to the finding that feedback, while helpful on average, makes performance worse in a large share of cases. In a 1996 meta-analysis of 607 effect sizes and 23,663 observations, Avraham Kluger and Angelo DeNisi found that feedback improved performance by about 0.41 standard deviations on average, but that more than a third of the effects were negative. The harm could not be explained by chance or by whether the feedback was positive or negative.
What is feedback intervention theory?
Feedback intervention theory is the explanation Kluger and DeNisi proposed for that pattern. It holds that feedback shifts attention between three levels: the details of the task, motivation to do the task, and the self. Feedback helps most when it keeps attention on the task and helps least when it pulls attention toward the self, for example toward what a score says about your ability or how you compare with others. The authors described the theory as preliminary, and it is best treated as a strong organising idea rather than a settled law.
Are grades or comments better for learning?
In primary and secondary school the evidence mostly favours comments. In Ruth Butler's 1988 study of 132 pupils in twelve classes, comments sustained interest and performance while grades did not, and giving a grade together with a comment worked like a grade alone. A 2019 set of meta-analyses found that grades led to worse achievement and motivation than comments. Newer studies add nuance: grade penalties can fade as students adapt, and grading systems that allow revision and reassessment have raised achievement.
Is negative feedback worse than positive feedback?
Not in a simple way. The sign of feedback did not predict which feedback backfired in the original meta-analysis. A 2018 review of 78 studies found that negative feedback did not lower intrinsic motivation compared with neutral or no feedback, and was less demotivating when it explained how to improve. Positive feedback can also backfire, for example on detail-oriented tasks or when it praises the person rather than the work.
Does praising a child for being smart really hurt?
The original 1998 studies found that children praised for intelligence persisted less and performed worse after failure than children praised for effort. However, a 2019 replication with 624 children aged 9 to 13 did not reproduce the effects on motivation, and a 2023 meta-analysis found growth mindset interventions had very small effects on achievement. The safer conclusion is that praise aimed at the work tends to be more useful than praise aimed at the person, while the size of the harm from intelligence praise remains uncertain.




