Introduction

There is a chart on a fridge somewhere with your handwriting on it. Twenty minutes of reading gets a sticker. Ten stickers gets a trip for ice cream. It is a reasonable thing to do. It is what the school suggested, what the library does over the summer, and what a pizza company built a forty-year marketing programme on. And you may have heard, somewhere, that it backfires.

That claim is one of the most repeated findings in popular psychology. Reward a child for reading and the child reads less.

It comes with a study attached, usually the one where preschoolers were promised a certificate for drawing and then stopped. It comes with a moral too. Stop bribing children.

Here is the problem with that story. It is half right, and the half that is missing is the half you can actually use.

The research it rests on is real and it has been replicated for fifty years. But the meta-analysis everybody quotes as proof, a 1999 paper covering 128 studies, reports in the same abstract that verbal praise increased people's willingness to keep going at an effect size of d = 0.33 [16]. It reports that rewards handed over as a surprise, with no promise beforehand, did nothing at all. The harm it found was specific and narrow. It belonged to rewards that were expected, tangible, and given for doing the thing.

So the sentence "rewards kill motivation" is not what the evidence says. What the evidence says is stranger and more useful.

Some rewards move the felt cause of an action out of the person and into the prize, and once that has happened, removing the prize removes the reason. Other rewards do the opposite. They tell a child they are good at something, and that is a message that survives the reward ending.

This article is about the theory that explains the difference, and about what it predicts when you put reading into the reward slot. It also covers the parts nobody mentions: a forty-year argument between two camps that was never settled, a study of the world's most famous reading-incentive programme that found nothing either way, and a small experiment with seventy-five third-graders that clears up most of the confusion.

Where the field agrees, this article says so. Where it does not, it names who is on each side.

Empty wooden reading table with open book and gold foil stars.

The Award That Emptied the Drawing Table

Start where the field started, because the experiment is small enough to hold in your head.

In 1973 Mark Lepper, David Greene and Richard Nisbett went into the Bing Nursery School at Stanford and watched which children chose to draw with felt-tip markers during free play [1]. They were not looking for children who needed encouragement. They were looking for the opposite. They wanted children who already loved it. The 51 children aged three and four who made the cut were each taken aside individually and put into one of three conditions [1]. In the first, the child was told beforehand that drawing would earn a Good Player Award, a certificate with a gold star and a red ribbon. In the second, the child drew with no mention of any award and was handed the same certificate afterwards as a surprise. In the third, no award was mentioned and none was given.

Then everyone went back to normal nursery school. A week or two later, observers counted how much of free play each child spent at the marker table.

Of those 51 children, the ones who had been promised the certificate spent about half the time drawing that the other two groups did [1]. The surprise group looked like the control group. Nothing separated them.

Read the design again, because the structure of it is the whole point. The only thing that differed between the promised group and the surprise group was whether the child knew about the award in advance.

The certificate was identical. The drawing time during the session was identical. What changed was the reason for drawing, and the reason turned out to be the durable part.

Two years earlier, Edward Deci had found the same thing in adults with a very different task [2]. In the first experiment of that 1971 paper, 24 students worked on SOMA puzzles, the kind that snap together into shapes [2]. Half were paid a dollar per puzzle solved. Then the experimenter left the room for eight minutes on a pretext, and what was measured was what the students did with the free time. The paid group went back to the puzzles less than the unpaid group. The same paper contained a third study that almost nobody mentions when they cite the first two. In that one the reward was not money. It was the experimenter saying, in effect, that was very good, you solved that faster than most people. Across the students in that condition, verbal praise did not reduce free-choice persistence [2]. It raised it.

Hold onto that.

It is the single most frequently deleted finding in this entire literature, and it has been sitting in the founding paper the whole time.

In 1975 Lepper and Greene showed with 80 preschool children that you do not even need a prize at all [3]. Those children worked on an activity while an adult watched them through a video monitor, and being watched was enough to produce the same later drop in voluntary engagement [3]. The undermining was not about receiving something. It was about the sense that the activity had become somebody else's business.

You have felt this.

A hobby that turned into a side income and stopped being fun. A book you loved until it appeared on a syllabus. The mechanism is not exotic.

Scattered colorful felt-tip markers on a pale wooden surface.

Three Needs, and a Theory That Is Not Really About Rewards

Self-determination theory did not begin as a theory of rewards. Deci and Richard Ryan named it in their 1985 book Intrinsic Motivation and Self-Determination in Human Behavior, published by Plenum Press, and the reward experiments were one line of evidence inside a much larger claim about what people need in order to keep going at anything.

The claim is that three psychological needs are basic, in the strong sense that satisfying them supports growth and thwarting them predicts decline [4].

Autonomy is the sense that what you are doing is yours. Not that you chose it from nothing, and not that nobody asked you to do it, but that you stand behind it. You can feel the difference between an evening you chose and one you were marched into.

Competence is the experience of being effective, of the thing working when you try. Relatedness is the sense of mattering to people who matter to you.

Autonomy is the one that gets misread, constantly. It does not mean independence and it does not mean the absence of structure. Ryan and Deci spent a 2020 review restating this because the misreading is so persistent [5]. A child can be told exactly what to do and still act autonomously if they endorse the reason. A child can be left entirely free and feel coerced by a peer group. Structure and autonomy support are separate dimensions and they combine perfectly well.

The theory also refuses the binary that most people bring to it. There is no clean split between intrinsic motivation over here and extrinsic motivation over there. There is a continuum, running from having no motivation at all, through several grades of doing something because you have to, to doing it because you value it, to doing it because it is absorbing in itself [6]. Joshua Howard and colleagues tested that ordering meta-analytically in 2017 and found the continuum structure held up [9].

That matters here more than it looks. A child who reads because she wants to know how the story ends and a child who reads because it is valuable to be a reader are not in the same place, but neither is failing. The child in trouble is the one reading only because the chart is on the fridge.

Applied to classrooms, the three needs have been the organising frame for four decades of work [7]. Deci, Robert Vallerand, Luc Pelletier and Ryan set out the educational version in 1991 [10], and Frédéric Guay's 2021 overview is a good short statement of where it stands now [11].

One refinement is worth knowing because it changes what you look for. Need satisfaction and need frustration are not two ends of one scale [8]. A child can have low autonomy support without anything actively going wrong. A child whose autonomy is actively thwarted, by surveillance or by pressure, is in a different state, and it is the second that predicts the sharpest decline. If you want the neuroscience of what a rewarding signal does inside a learning brain, we have covered dopamine and learning separately.

Not All Extrinsic Motivation Is the Same

Before the rewards, one more piece of the theory, because without it the practical advice at the end of this article sounds contradictory.

Most people carry a two-box model. The good box is doing things you love, the bad box is doing things because somebody makes you. Self-determination theory does not use that model, and the reason matters.

Between pure enjoyment and pure coercion sit several distinct ways of doing something for a reason outside the activity itself, and they are not equally fragile [6]. At one end is doing it because there will be consequences if you do not. Slightly further along is doing it to avoid feeling guilty or to feel like a good person, which is still pressure, only now the pressure is internal and you are applying it yourself. Further along again is doing it because you have decided it is worth doing. Further still is doing it because it fits who you take yourself to be. Only at the far end is doing it because the activity is absorbing in itself.

Those are all real states, they can be measured separately, and the ordering has been tested. Howard, Marylène Gagné and Julien Bureau examined whether the sequence holds up as a continuum in 2017 and found that it does [9]. Adjacent forms correlate more strongly than distant ones, which is what a continuum predicts and what two separate boxes do not.

Apply it to a nine-year-old and the picture gets more useful straight away. A child who reads because she will lose screen time otherwise is in the most brittle state there is.

Remove the threat and the reading stops the same evening. A child who reads because reading is something she does, because she is a person who has a book on the go, is nowhere near that. The activity is attached to her sense of herself, and nothing has to be enforced.

The second child is not intrinsically motivated in the technical sense. She may not find every book gripping. She is doing it for a reason outside the pleasure of the moment. But the reason is hers, and that turns out to be almost the whole difference.

This is what the theory means by internalisation, and it is a process rather than a switch [11]. Reasons start outside a person and, under the right conditions, move inward. Under the wrong conditions they stay outside and need constant topping up, which is exactly what a sticker chart feels like by month three.

So the goal is not to get a child from extrinsic to intrinsic. That framing sets an unreachable target, because plenty of necessary reading is never going to be thrilling. The goal is to move the reason inward, and the conditions that do that are the same three needs.

Every Reward Says Two Things at Once

Here is the part of the theory that does the work, and it is only one idea.

Any reward carries two messages simultaneously. One is controlling. It says you are doing this because I am giving you something for it. The other is informational. It says you did that well. Which message dominates decides which way motivation moves [12]. Richard Ryan demonstrated the split experimentally in 1982 by taking the same reward and administering it two different ways, once with controlling language and once with informational language, and getting opposite effects on later engagement [12]. The object was constant. The framing was not.

That is why a gold star is not one thing.

A gold star handed over with the words you finally did what I asked is a very different object from the same star handed over with the words you got all the way through a chapter book this week.

Controlling

Informational

A reward is given

Which message lands?

Cause moves outward

Sense of competence rises

Interest drops later

Interest holds or grows

Ryan, Valerie Mims and Richard Koestner formalised the categories in 1983, and their typology is the one the entire later literature runs on [13]. A reward can be given regardless of what you do, given for doing the task at all, or given for doing it to a standard. Those three are not variations on a theme. They behave differently, and the differences are large enough to change what you should do on a Tuesday evening.

There is also a route in the other direction, and it is the reason this article does not end with never reward anything. In 1994 Deci, Haleh Eghrari, Brian Patrick and Dean Leone showed that a genuinely dull task can be taken up willingly if three things happen together: a real reason is given, the person's reluctance is acknowledged out loud, and controlling language is avoided [14]. Not one of those three is a prize.

Internalisation is the technical name for that. It is what you are hoping for when you want a child to read without being asked.

The Argument You Were Never Told About

Now the part that gets left out of every explainer on the first page of any search engine.

In 1994 Judy Cameron and W. David Pierce published a meta-analysis in Review of Educational Research covering 96 experimental studies that compared rewarded participants against unrewarded controls [15]. Their conclusion was close to the opposite of the received wisdom. Overall, they reported, reward does not decrease intrinsic motivation. Verbal praise increases it, and their own figure for that was d = 0.38 on free-choice behaviour [15]. The only negative effect they found was confined to expected tangible rewards given for simply doing a task, which came out at d = -0.21, and they described it as minimal [15].

They did not stop at a null result. They wrote that there was no reason not to use reward systems in education, and they called for abandoning cognitive evaluation theory outright [15]. That is on page 396 of the paper.

This was not a fringe publication. Review of Educational Research is one of the field's most visible journals, and for years the 1994 paper was what people cited to say the reward panic was overblown.

Five years later Deci, Koestner and Ryan replied in Psychological Bulletin with a meta-analysis of their own, covering 128 studies drawn from 94 published articles and 19 dissertations [16]. Of those, 101 included a measure of free-choice behaviour and 84 included a self-report measure of interest. Their argument was methodological rather than rhetorical. They said the 1994 analysis had collapsed across categories where the interactions were theoretically meaningful, had used inappropriate control groups, and had discarded close to a fifth of the studies as outliers rather than asking why those studies behaved differently.

Rerun with the categories kept apart, the undermining effect came back.

The same 1999 issue carried two replies. Robert Eisenberger, Pierce and Cameron argued that the effects of reward were negative, neutral and positive depending on conditions, and that the negative case had been overgeneralised [17]. Lepper, Jennifer Henderlong and Isabelle Gingras came at the meta-analytic method itself, under a title about the uses and abuses of meta-analysis [18].

Two papers arguing in opposite directions inside one issue of one journal is unusual, and it tells you how much was at stake.

Reward systems were, and are, the operating assumption of most schools.

It did not end there. In 2001 Deci, Koestner and Ryan published a paper called Extrinsic Rewards and Intrinsic Motivation in Education: Reconsidered Once Again [19]. In the same year Cameron, Katherine Banko and Pierce published one called Pervasive Negative Effects of Rewards on Intrinsic Motivation: The Myth Continues [20]. Those two titles, published in the same year, are the state of the field.

So what is the honest summary?

The mainstream of educational and motivational psychology accepts the 1999 analysis. Most textbooks teach it. The behaviour-analytic tradition never conceded, and its objection was never simply stubbornness. Cameron and Pierce were wrong about the method, and they were right that never reward anything is not a conclusion the data support.

One more voice belongs here, because it is why the strong claim reached the public at all. Alfie Kohn's 1993 book Punished by Rewards, published by Houghton Mifflin, took the laboratory finding and turned it into an argument against gold stars, grades, and pay for performance.

It sold widely and it shaped a generation of school policy. It is also considerably more certain than the evidence underneath it.

What the Meta-Analysis Actually Says

Set the argument aside and look at the numbers yourself, because they are more interesting than either camp's headline and you will not see them quoted often.

These are the effects on free-choice behaviour from the 1999 meta-analysis, which is to say what people voluntarily did once nobody was making them [16].

What you giveEffect on voluntary persistenceStudies behind it
Reward for engaging with the taskd = -0.4055
Reward for completing the taskd = -0.3620
Reward for performing to a standardd = -0.28not separately reported here
All expected tangible rewardsd = -0.3692
Unexpected tangible rewardd = +0.019
Verbal praise and positive feedbackd = +0.3321

Look at the bottom two rows before you look at anything else. They are the whole practical payload of fifty years of research and they are almost never quoted.

A reward you did not know was coming has an effect of essentially zero. Not small. Zero, at d = 0.01 across nine studies [16]. There is nothing to undermine, because there was never a deal. Praise runs the other way at d = 0.33 on behaviour and d = 0.31 on self-reported interest [16]. Telling a child something true about what they did is not a reward in the sense that damages anything. It is information, and information about your own competence is one of the three things the theory says you need.

The self-report numbers are also worth a moment because they are smaller than the behavioural ones. Rewarding engagement moved self-reported interest by d = -0.15 and rewarding completion by d = -0.17 [16]. What people said about liking the task moved less than what they chose to do. The damage shows up in behaviour before it shows up in feeling, which means a reward scheme can be hollowing something out while everyone involved still reports enjoying it.

Children Take More of It Than Adults

The title of this article makes a claim about children specifically, and that is not decoration. It is in the data.

Across 57 free-choice studies involving children, tangible rewards produced an effect of d = -0.39, and the difference between children and college students was itself statistically significant [16]. Tangible rewards hurt children more. Verbal praise went the other way on the same comparison. Across 21 studies, praise raised free-choice persistence overall at d = 0.33, but the effect for children alone was d = 0.11 and did not reach significance [16]. Children get more of the harm from prizes and less of the benefit from praise than adults do.

Why would that be? One explanation offered inside the literature is developmental. Young children build concrete schemas about what adults use incentives for, and the schema is roughly this: grown-ups pay you to do things nobody would do for free. A five-year-old offered a prize for reading has just been given evidence about reading.

There is a 1987 experiment by Wendy Grolnick and Ryan that puts this uncomfortably close to our subject [21]. Children in that study were given a passage to read. Some were told they would be tested on it and should try to remember as much as possible, which is controlling framing. Others were introduced to it without pressure. The controlled group showed lower interest afterwards, and their rote retention decayed faster.

Reading was the task. Not puzzles, not markers. Reading, which is what you are here about.

Tall library aisle with wooden shelves and warm lighting.

Twenty-Eight People and a Stopwatch

For thirty years the undermining effect was a behavioural finding with a cognitive story attached. In 2010 Kou Murayama and colleagues took the effect into a scanner [22]. The task they built is almost comically simple. A stopwatch starts on its own, and you press a button trying to stop it within 50 milliseconds of the five-second mark. The study ran 28 right-handed volunteers [22]. The 14 participants in one group earned money for accurate presses and the 14 in the other did not [22]. Then the money stopped, and the measure became how often people voluntarily went back to the stopwatch during a free period.

The reward group returned less often. That is the behavioural undermining effect, reproduced in adults on a task nobody has any prior feeling about. What the imaging added was where it showed up. Activity in the anterior striatum and in prefrontal areas fell along with the behavioural drop [22]. The authors read this as the valuation system integrating the extrinsic reward value with the intrinsic task value, and the intrinsic side losing.

Be careful about what an image like that establishes. A drop in striatal activity is not a picture of motivation leaving a person.

It is a change in a signal that tracks a dozen other things, measured in a small group and averaged across trials. It is consistent with the behavioural result rather than independent confirmation of it, because the same participants produced both.

Stefano Di Domenico and Ryan reviewed this line of work in 2017 and made a point worth carrying forward: the neuroscience of intrinsic motivation is young, and the findings that exist are mostly small-sample imaging studies [23]. A scanner does not make a small sample bigger. That caution is not hypothetical. In 2026 a group led by Hiroaki Ayabe published a paper arguing more or less the opposite reading of the same phenomenon, under the title Beyond the undermining effect: extrinsic rewards preserve neural intrinsic reward [24]. Their case is that the intrinsic reward signal is not erased by extrinsic reward, and that what changes is something else. This is live and it is recent. Anyone telling you the brain basis of the undermining effect is settled has not read the last twelve months.

The positive-side mechanism is on firmer ground. In 2014 Matthias Gruber, Bernard Gelman and Charan Ranganath showed that being in a state of curiosity improved memory for material, including incidental material that had nothing to do with what the person was curious about, and that the effect ran through dopaminergic circuitry and the hippocampus [25]. Wanting to know something is not a soft outcome. It changes what gets encoded.

That connects to why a well-pitched difficulty level feels the way it does. We have written about flow states and optimal difficulty, which is the competence need in its most recognisable form.

Does Any of This Survive Outside the Laboratory?

A fair objection at this point is that a preschooler with a marker and a college student with a stopwatch are not a child with a reading log. So does the effect generalise?

Partly, and the qualifications are the interesting bit.

In 2014 Christopher Cerasoli, Jessica Nicklin and Michael Ford pulled together forty years of data and found something that resolves a lot of arguing [26]. Incentives and intrinsic motivation are not substitutes. Incentives predict how much of a thing gets done. Intrinsic motivation predicts how well. When the incentive is tightly tied to the performance, intrinsic motivation matters less for output. When it is loosely tied, intrinsic motivation is a strong predictor.

Apply that to reading and it lands immediately.

A points programme will increase the number of books logged. That is the quantity channel and it works. Whether it increases anything you would call reading is a different question with a different answer.

Marianne Promberger and Theresa Marteau asked in 2013 when financial incentives reduce intrinsic motivation, comparing the behaviours studied in laboratories against health behaviours in the field, and concluded that the conditions are narrower than the popular version implies [27]. The effect needs pre-existing interest to damage. For a behaviour nobody wanted to do anyway, there is nothing there to spend.

The workplace literature has been going through the same argument. A 2022 study framed extrinsic rewards for creativity in self-determination terms and found the outcome depended on how the reward was administered rather than on whether it existed [28]. And in 2025 Marylène Gagné and colleagues published work on why bonuses promote deviant behaviour, which is a different failure mode again [29]. A contingency does not only change how hard people try. It changes what they optimise for.

Keep that sentence. It comes back when we get to points-per-book systems.

1971
Edward Deci pays 24 students to solve puzzles and they play with them less afterwards
1973
Lepper Greene and Nisbett promise 51 preschoolers a certificate for drawing
1975
Lepper and Greene show that being watched produces the same drop with no prize at all
1982
Richard Ryan separates the controlling and informational sides of the same reward
1985
Deci and Ryan name self-determination theory in a book
1993
Alfie Kohn takes the finding to the public in Punished by Rewards
1994
Cameron and Pierce meta-analyse 96 studies and reject cognitive evaluation theory
1999
Deci Koestner and Ryan answer with 128 studies and two replies appear in the same issue
2001
Both camps publish again and neither concedes
2006
A randomised trial of Accelerated Reader finds achievement gains in 978 students
2008
Marinak and Gambrell put a book in the reward slot and the undermining disappears
2010
Kou Murayama images the striatum going quiet after the money stops
2024
A synthesis of 637 samples and 388.912 people maps need support and need thwarting
2026
A new imaging paper argues the intrinsic reward signal is preserved after all

Fifty-five years, and the disagreement in the middle of it has never closed. You are reading a live argument, not a settled one.

Now Put Reading in the Slot

Reading motivation has its own research literature, largely separate from the reward experiments, and it is worth meeting on its own terms before the two are combined.

The modern version starts with Allan Wigfield and John Guthrie, who built a multidimensional questionnaire in 1997 and used it to relate motivation to how much and how broadly children read [30]. Reading motivation turned out not to be one thing. It splits into curiosity, involvement, challenge, competition, recognition, grades, compliance and several more. Two of those clusters matter here. Intrinsic reading motivation is reading because you want to be in the book. Extrinsic reading motivation is reading for grades, recognition, or because you were told to. They are measurable separately and they behave differently [32]. Samantha Ives and colleagues reviewed the measurement question in 2022 and their conclusion is a useful caution: the field has not fully agreed on what it is measuring, and studies that look contradictory are sometimes using different instruments [33].

Now the pattern, which is consistent enough across two decades that it is worth stating plainly, and consistent enough that you can hold it in one sentence.

In 2004 Judy Wang and Guthrie modelled intrinsic motivation, extrinsic motivation, amount of reading and past achievement together in American and Chinese samples [31]. Intrinsic motivation predicted text comprehension, and the path ran through how much children read. Extrinsic motivation did not behave the same way.

Notice what is carrying the effect there. Not motivation acting directly on comprehension, which would be a strange thing to claim, but motivation acting on how much reading happens, and reading amount acting on comprehension.

That chain is unglamorous and it is the most durable finding in this whole literature. Wanting to read gets you to the page. The page does the rest.

The strongest single result comes from Germany. Michael Becker, Nele McElvany and Marthe Kortenbruck followed 740 students from grade 3 into grades 4 and 6 [34]. Intrinsic reading motivation predicted later reading literacy, mediated by reading amount. Extrinsic reading motivation was a negative longitudinal predictor of later literacy, and it stayed negative when reading amount and earlier literacy were controlled. That is not a correlation between two survey scales. Across 740 students it is a child's reason for reading in grade 3 predicting how well they read three years later, in the wrong direction [34].

A negative predictor is a strong claim and it is worth stating carefully. It does not mean that being told to read makes a child worse at reading.

It means that among children matched on how much they read and on how well they already read, the ones whose stated reasons were external did worse later. Something about the reason itself carried information the reading volume did not.

Ellen Schaffner, Ulrich Schiefele and Hannah Ulferts confirmed the mediation route in 2013, with reading amount doing the work between motivation and comprehension [35]. Margaret Troyer and colleagues found the same shape in 2018 [36]. At the top of the evidence pyramid sits a 2020 meta-analysis by Jessica Toste and colleagues covering kindergarten through twelfth grade [37]. It pulled 132 articles, 185 independent samples and 1,154 effect sizes, and found a moderate overall relation between motivation and reading achievement at r = .22. Which motivation construct was measured moderated the size of that relation. Reading domain did not.

An r of about .22 is a modest number and deserves an honest translation. It is not nothing. It is also not the kind of relationship that predicts one child's reading from their motivation questionnaire.

Across a class it shows up. Across a single kitchen table it is swamped by everything else, which is worth remembering before anyone redesigns a family evening around it.

Volume is doing a great deal of the work in all of this, and there is a separate meta-analysis on that. Suzanne Mol and Adriana Bus analysed 99 studies covering 7,669 participants and found print exposure related to comprehension and technical reading across every age band from preschool to university [38]. Children who read more get better, and children who are better read more.

Two results sharpen the picture rather than repeating it. Sarah Logan, Emma Medford and Naomi Hughes reported in 2011 that intrinsic motivation mattered most for the weakest readers, which is the opposite of treating motivation as a luxury for children who already cope [39]. Jessie De Naeghel and colleagues split recreational from academic reading motivation in 2012 and found the recreational side, not the school-facing side, carrying the relation to engagement and comprehension [40].

There is a pattern to how these studies are built and it limits what any of them can say. Children fill in a questionnaire about why they read. Later, someone measures how well they read.

The reasons children give for their own behaviour are not always the reasons operating. That does not make the work useless. It makes it a measure of stated orientation rather than of motivation.

Two studies complicate the simple story in a useful way. Ai Miyamoto, Maximilian Pfost and Cordula Artelt reported reciprocal relations in 2017, with motivation and competence each feeding the other [41]. Karin Hebbecker, Natalie Förster and Elmar Souvignier found the same reciprocity in 2019 [42]. So the arrow is not clean, and anyone drawing it as one-directional is simplifying.

Neither of those results is a footnote. Reciprocity changes what an intervention can be expected to do. If motivation and skill each feed the other, then a push on either side propagates, and a child stuck at the bottom is stuck in a loop rather than missing a single ingredient. That is hopeful, because there are two places to push instead of one. It is also harder, because a loop resists a single push in a way a chain does not.

More recent work keeps replicating the core relation in different populations. Chinese adolescents in 2020 and Hong Kong students in 2021 produced the same mediation shape, with reading amount and reading strategy carrying the effect from motivation through to achievement [43] [44]. Two school systems, one mechanism.

Two European samples asked slightly sharper versions of the question. A 2021 study of Flemish children asked whether skill or will contributes more to comprehension, and found both mattering rather than one winning [45]. A 2022 Swedish study measured how reading felt to students alongside how much of it they did, treating the two as separate outcomes, which is the distinction most reward programmes collapse [46].

One newer approach is worth the space because it changes what the averages mean. Instead of treating intrinsic and extrinsic motivation as two scores to correlate, profile studies ask which combinations actually occur in real children. A 2022 study of Chinese adolescents found distinct profiles rather than one continuum, and those profiles related differently to how much the children read [47]. That matters for a reward scheme. A child high on both kinds of motivation is not the same case as one high only on the external kind, and the average of the two describes neither.

Two findings complicate any picture built on print alone. A 2022 study of digital reading practices found motivational patterns print research had not described, so a child who reads constantly on a screen may not be measured properly by the instruments above [48]. And a 2023 analysis by Margriet van Hek and Gerbert Kraaykamp asked why girls read more than boys and found that parents and schools stimulate the two differently, which relocates part of the gap from the child to the environment [49].

None of these studies is an experiment on rewards.

They are the background against which a reward programme gets introduced, and they tell you what the target actually is. Not books logged. Reading amount, sustained, driven by wanting to.

The Pizza That Proved Nothing

If reward programmes for reading were as damaging as the strong claim implies, the most famous one on earth should have left a mark.

Book It! was launched by Pizza Hut in 1984. Children read a set number of books in a month, a teacher signs off, the child gets a certificate for a personal pan pizza.

It has run for four decades and passed through tens of millions of American children. It is exactly the design the laboratory work says should backfire: tangible, promised in advance, contingent on finishing.

In 1999 Stephen Flora and David Flora went looking for the damage [50]. They surveyed roughly 107 students about whether they had taken part in Book It! as children, whether they had ever been paid money for reading, how much they read now, and how much they enjoyed it [50].

They found nothing. Neither participation in Book It! nor having been paid for reading was associated with reading more or reading less as an adult, and neither was associated with intrinsic motivation for reading [50]. The self-report answers were mixed in the way self-report answers usually are. About 28 percent of those 107 students said the programme increased their enjoyment of reading and about 64 percent said it made no difference [50].

Be careful with what this study can and cannot carry. It is retrospective, it is self-report, and the 107 participants are college students, which is a group selected for having survived school reading reasonably well [50]. It cannot tell you what happened to children who dropped out of the pipeline, it has no control over how individual teachers ran the programme, and a null in a sample that size is not proof of absence.

But it is the evidence that exists, and the evidence that exists says the most-cited example of reward damage did not produce measurable reward damage.

That result is inconvenient for the argument here, which is why it gets a section rather than a footnote. The laboratory effect is real and heavily replicated. The most obvious field application of it shows nothing. Something in between is doing work.

Well-worn hardback books stacked on wooden floor with soft shadows.

Accelerated Reader, and the Outcome Nobody Measured

The other giant is Accelerated Reader. A child reads from a levelled library, takes a short computerised quiz, and earns points scaled to the book's difficulty. Points accumulate, and many schools attach prizes to them.

In 2006 John Nunnery, Steven Ross and Aaron McDonald ran the trial that people should cite and mostly do not [51]. Teachers in grades 3 to 6 were randomly assigned either to implement the programme or to serve as controls, covering 978 urban students [51]. Reading achievement was measured with a three-level model of growth trajectories. Students in the programme classrooms grew faster. Across those 978 students the trial reported effect sizes of +0.07 to +0.34 depending on grade [51]. That is a real randomised result and it should not be waved away by anyone who dislikes points systems.

Now read the outcome variable again.

Reading achievement. Not intrinsic motivation, not voluntary reading five years later, not whether these children became adults who read. The trial measured the thing the programme was sold on and found it delivered.

That is not a criticism of the study. It is what a randomised trial of a reading programme usually measures, and it is why the two literatures talk past each other.

The reward researchers measure what people do when nobody is watching. The programme evaluators measure test scores at the end of the year. Both are legitimate. They are not the same outcome and they do not have to move together.

Programme or studyDesignWhat was measuredResult
Book It!retrospective survey of about 107 college studentsreading amount and intrinsic motivation years laterno effect in either direction
Accelerated Readerrandomised trial with 978 students in grades 3 to 6reading achievement growthfaster growth, +0.07 to +0.34
Book versus token as the rewardrandomised experiment with 75 third-graderswillingness to keep reading afterwardsbook and no-reward beat token

The third row is the one you want.

A Book Is a Reward Too

In 2008 Barbara Marinak and Linda Gambrell ran the experiment that should be far better known than it is [52]. A total of 75 children in third grade were randomly assigned to one of five groups [52]. Four of those groups received a reward for reading, in a two-by-two design: the reward was either a book or a token, and the child either did or did not get to choose it. The fifth group was a control that received nothing at all. Afterwards, the researchers measured task persistence, which is to say whether children went back to reading when they had the option.

The design is the clever part.

Every child in the four reward conditions got something, so the comparison is not reward against no reward. It is one kind of reward against another.

Among those 75 children, the ones who received a book scored like the ones who received nothing [52]. The children who received a token scored worse [52].

Read that again with the theory in hand and it stops being surprising.

A token is a controlling message. It says reading is the price of the token, and once the token is in your hand the price has been paid.

A book is not outside the activity. It is more of the activity. Getting one is closer to competence feedback than to payment, and it leaves you holding something you can read.

Marinak and Gambrell called this the proximity of the reward to the desired behaviour. Lorilynn Brandt and colleagues developed it for practitioners in 2024 under the heading of proximal rewards [53]. Gambrell had already set out the broader engagement principles in 2011 [54], and access to books and choice among them sit near the top of that list.

Be careful how much this one study is asked to reconcile. The proximity idea predicts that Book It!, which paid in pizza, should have done the most damage of any programme here. It did not. The one follow-up found nothing either way [50], and a null is a problem for the prediction rather than support for it. The Accelerated Reader trial does not rescue it, because that trial measured achievement and this argument is about motivation [51]. Those are the two outcomes this article has already insisted on keeping apart.

So the honest version is narrower. Marinak and Gambrell is the only study here that changed the reward while holding everything else still, and inside that design the reward type decided the result. That is a finding about one controlled comparison, not a theory of every programme ever run.

Inside that comparison, the variable that mattered was not whether a reward was given.

It was what the reward said about reading.

Note the sample size before you build a policy on it. That is 75 children in one design and one age group [52]. It is a small study and it deserves replication rather than reverence. But it is the only experiment in this area that varied the reward type directly, which is precisely the variable the theory says should matter.

No, given afterwards

Yes, promised

A book, a library trip

A token, food, screen time

Rewarding reading

Promised in advance?

Little risk to interest

Reward is reading?

Inside the activity

Outside the activity

Reason leaves too

There is a second-order effect in points systems that deserves naming, because it is the reading version of what happens when bonuses reshape behaviour at work [29]. When points scale with book difficulty, the optimal strategy is not to read what you want. It is to read whatever maximises points per unit of effort. Teachers report this pattern in strong readers, who learn to hunt the point-efficient title rather than the interesting one. That is a contingency doing what contingencies do, separate from anything the child feels about reading. If you want the case for effort that feels unpleasant and works anyway, we have covered desirable difficulties elsewhere.

What Thwarts Each Need at a Reading Table

The three needs are abstract until you put them somewhere specific. Here is what supporting and thwarting each of them looks like when the activity is reading.

NeedWhat supports itWhat thwarts it
Autonomychoice of book, choice of when, a reason given rather than an orderassigned titles only, a nightly log countersigned by an adult, reading as a condition of something else
Competencea book at the right stretch, feedback about what improvedpublic levelling, quizzes that punish an ambitious choice, comparison with a sibling
Relatednessadults who visibly read, being read to past the age it seems necessary, talking about booksreading as solitary compliance, no adult reader anywhere in the house

This is not a list of tips for you to work through. It is a map of where things go wrong, and the evidence behind it is substantial.

Pedro Conesa and colleagues reviewed how the three needs behave in elementary and middle school classrooms specifically in 2022 [55]. The largest evidence base is a 2024 meta-analysis by Howard and colleagues that synthesised 8,693 correlations from 637 samples covering 388,912 students [56]. It examined six categories of need-supportive and need-thwarting behaviour by teachers and parents. Supportive behaviour correlated with performance, engagement and wellbeing. Thwarting behaviour ran the other way.

The thwarting side has its own literature and it is bleaker than the supportive side is cheerful. Bart Soenens and Maarten Vansteenkiste rebuilt the concept of parental psychological control in 2010, separating pressure that is applied to behaviour from pressure applied to the child's internal states [57]. Guilt, conditional approval and love withdrawal are the tools of the second kind, and they are the ones that do lasting damage.

A reading chart is not psychological control. But a reading chart that becomes a nightly negotiation about whether you are disappointed in someone can drift toward it, and that drift is worth watching for. Reading is also social, and children copy what they see; we have written about what social learning and the Bobo doll experiment actually established about modelling.

What Works Instead

If the reward is not the lever, what is?

The clearest answer from the last five years is need-supportive teaching, and the reading-specific evidence has become quite good.

Joseph Haw and Ronnel King showed in 2022 that need-supportive teaching predicts reading achievement, and that the path runs through intrinsic motivation rather than around it [58]. Support the three needs, motivation rises, achievement follows. Remove the middle step and the relation weakens.

That middle step is the part worth holding onto. It says the teaching does not raise scores by drilling. It raises them by changing what a child wants, and the wanting then produces the reading that produces the scores. Take out the wanting and you are back to a points system, which raises the count and leaves the rest alone.

There is a depressing regularity in the reading literature that makes this urgent. Intrinsic reading motivation declines across school years, in almost every sample anyone has measured. Laura Engler and Andrea Westphal reported in 2024 that teacher autonomy support counters that decline [59]. It does not merely correlate with higher motivation at one moment. It changes the slope.

A slope is a better thing to move than a level.

Raising a child's motivation for a term is worth little if the trend underneath it keeps falling. Bending the trend is what actually compounds.

Anna Hawrot and Ji Zhou found in 2023 that changes in how students perceived their teacher's behaviour predicted changes in their intrinsic reading motivation, which is the within-person version of the same claim [60]. A 2025 study reported reciprocal effects among autonomy support, intrinsic motivation and reading achievement, each feeding the next [61]. Some of the effective moves are startlingly small. In 2015 Ivar Bråten, Roy-Petter Johansen and Helge Strømsø varied only how a reading task was introduced, and got differences in both intrinsic motivation and comprehension [62]. Same text. Same students. Different sentence beforehand.

Think about the sentence you say before handing a child a book. Not a threat, not a promise, and not nothing. A reason.

Here is why this one, here is what you might find in it, here is what I thought when I read it. That takes fifteen seconds and it is doing more work than a month of stickers.

Under all of this sits a developmental story about interest itself. Suzanne Hidi and Ann Renninger proposed in 2006 that interest develops in four phases, beginning with situational interest triggered by something in the environment and, if supported, becoming a stable individual interest that a person maintains on their own [63]. That model is the alternative to the reward model, and it explains why a single gripping book can matter more than a term of stickers. Guthrie and colleagues took the practical version of this into classrooms and reported motivation and comprehension growth together in the later elementary years [64]. James Kim and colleagues showed in 2021 that a content-literacy intervention could raise comprehension, domain knowledge and reading engagement at the same time [65]. Those three outcomes usually get traded off against each other. They do not have to be.

The four-phase model is also the most reassuring idea in this article, because it says the starting point does not have to be love. It can be a single triggered moment of curiosity, provided somebody notices it and feeds it.

A child who is gripped by one book about sharks is not yet a reader. She is at phase one, and phase one is where every reader started.

The teaching style itself is learnable, which is the most useful fact in this section. Johnmarshall Reeve and Sung Hyeon Cheon argued in 2021 that autonomy-supportive teaching is malleable and that training it produces measurable gains [66]. In 2024 Reeve and Cheon reported that the training works by teaching perspective taking first, with the classroom practices following from it [67]. In the same year Cheon and colleagues found the effect transferring from teachers to parents [68].

A 2024 meta-analysis of self-determination theory interventions in education pooled 137 effect sizes across 9,433 participants [69]. It found consistent support for the autonomy and competence needs, and partial effects on intrinsic motivation itself.

Partial is the honest word. These interventions reliably change how supported children feel. They change what children want less reliably, and what children do less reliably still. Anyone selling a programme on the strength of this literature is selling the first link in a chain and charging for the last.

Even the framing of a single task can be manipulated. A 2023 study varied whether instructions for an online learning task were need-supportive and measured the difference in intrinsic motivation [70]. Julien Bureau and colleagues meta-analysed the antecedents of autonomous and controlled motivation in 2021 and found the pattern holding across a large body of school research [71].

None of this transfers automatically from a classroom to a living room. A teacher has thirty children and a curriculum. A parent has one child and a bedtime. What carries across is the stance rather than the technique: treating a child as somebody with reasons rather than somebody to be arranged.

At home, the same variables show up under different names. Grolnick and Ryan reported in 1989 that parents who supported autonomy had children with better self-regulation and school competence [72]. Mireille Joussemet, Renée Landry and Koestner set out what a self-determination approach to parenting means in practice in 2008 [73]. A 2015 meta-analysis by Ariana Vasquez and colleagues found parental autonomy support related to both achievement and psychosocial functioning [74].

Two studies answer the obvious objection, that this is just a nicer way to describe children who were doing well anyway. Hyungshim Jang, Reeve, Ryan and Ahyoung Kim asked Korean students in 2009 what made a lesson productive and satisfying, and the answers came back as autonomy, competence and relatedness rather than anything resembling an incentive [75]. Geneviève Taylor and colleagues then tracked school achievement over time in 2014 and found that it was the autonomous forms of motivation, not the controlled ones, that predicted where a student ended up [76].

None of that requires a chart on a fridge. Some of it requires a chair, a lamp, and an adult who is visibly reading something. If you want to know how a behaviour becomes automatic without external pressure, we have covered how habits become automatic separately, and the long-horizon version of the same question in the psychology of expertise.

Where the Evidence Runs Out

An article like this should end by saying what it does not know. There is more of that than the confident version admits.

The first gap is causal direction, and it is a large one. Almost all of the reading motivation literature is correlational or longitudinal-correlational. It shows that intrinsic motivation and reading achievement travel together and that motivation earlier predicts achievement later. It does not show that raising motivation raises achievement.

In 2022 Elsje van Bergen and colleagues attacked the assumption head on, using a twin design, in a paper titled plainly enough that literacy skills seem to fuel literacy enjoyment rather than the other way around [77]. If that is right, then a substantial part of the field has its arrow backwards, and children enjoy reading because they got good at it rather than getting good because they enjoyed it. That does not overturn the reward experiments, which are randomised and which measure something else. But it should make you cautious about any claim that fixing motivation fixes reading. The reciprocal findings from 2017 and 2019 suggest the truth is that both directions operate [41], and a mechanism that runs both ways is much harder to intervene on than one that runs in a line.

Twin designs are good at this question and bad at others. They separate what runs in families from what does not, which is what you need when skill and enjoyment travel together.

They cannot tell you what would happen if you changed a child's environment, because that is not the comparison they make. So take the direction seriously and hold the policy conclusion loosely.

The second gap is duration. The undermining experiments measure free-choice behaviour minutes or days after the reward stops. Almost none of them follow anyone for years. The one study that did look years later, at Book It!, found nothing [50]. We do not have good evidence about what a childhood of sticker charts does to a forty-year-old, because that study has not been done.

The third gap is the one running through this whole article. The core dispute from 1994 to 2001 was never resolved by new data. It was resolved by one camp becoming the mainstream and the other camp continuing to publish. Deci, Koestner and Ryan's methodological criticisms of the 1994 meta-analysis were substantial and largely accepted [16]. Cameron, Banko and Pierce were still calling the pervasive negative effect a myth in 2001 [20]. Nobody ran a large pre-registered replication to settle it, and the neuroscience that might have adjudicated is itself now contested [24].

The fourth gap is that most of this research was done on children in North America and Western Europe, in school systems with particular assumptions about reading. The cross-cultural work mostly replicates the pattern, but replicating it in a second wealthy country is a modest test.

So the honest position is this. The laboratory mechanism is real, well replicated, and strongest in exactly the population the title is about.

The field results on actual reading programmes are mixed, and the one variable that consistently explains the mix is what the reward has to do with reading. And the confident version you will find on most pages about this topic is repeating a conclusion that its own founding meta-analysis qualifies in the abstract.

What This Changes About the Chart on the Fridge

Put the evidence in one place and the practical shape of it is clearer than the argument around it suggests.

The founding demonstration is small and it is old. It rests on 51 children at a nursery school in 1973, of whom only a third were in the condition that mattered, and on 24 students solving puzzles in 1971 [1]. Those are not large studies and nobody should pretend otherwise. What makes the finding load-bearing is that it kept reproducing. By 1999 there were 128 studies to pool, and the effect survived being pooled [16]. The rival 96-study analysis drew on much of the same literature and read it differently, so the two pools are not 224 separate demonstrations.

Two small founding studies are a thin foundation for a claim this widely repeated, and it is fair to say so plainly. What rescues it is not the original experiments.

It is that the effect kept turning up when other people went looking, in other laboratories, with other tasks and other ages, for five decades. Replication is doing the work that sample size cannot.

The size of the effect is moderate rather than dramatic. Rewarding a child for engaging with an activity moved voluntary persistence by d = -0.40, and rewarding completion by d = -0.36 [16]. In the same analysis, across the studies that used children rather than adults, tangible rewards came out at d = -0.39 while praise for those children came out at d = 0.11 and did not reach significance [16]. Those are the numbers behind the headline, and they describe a real effect that is not a catastrophe.

Moderate effects get misreported in both directions. A headline turns d = -0.40 into rewards destroy motivation. A sceptic turns the same number into a laboratory curiosity. Neither is right. An effect that size, across a school year, is real and also not the largest thing happening in that school.

Then look at what happened when anyone took it into a school. In roughly 107 students surveyed years afterwards, the most famous reading-reward programme in the world left no trace in either direction [50]. In a randomised trial covering 978 students, a points-based programme raised reading achievement with effect sizes of +0.07 to +0.34 [51]. And in the one experiment that changed the reward itself rather than the programme around it, 75 children divided across five conditions came out better with a book than with a token [52].

None of those three studies is decisive on its own. Together they point somewhere specific, and it is not where the popular version points. The question was never how much reward to use. It was what the reward is made of.

There is one more asymmetry worth carrying away. The effect on what people say about an activity is smaller than the effect on what they choose to do, at d = -0.15 for engagement-contingent rewards against d = -0.40 on behaviour [16]. That gap is the reason this is hard to notice from inside a family or a classroom. Everyone will tell you the reading is going fine. The change shows up on the evenings when nobody suggests it.

If you take one thing from all of it, make it this. The question is not whether to reward. It is what the reward tells a child about why they are reading, and that is the variable almost nobody controls. A book says keep going. A token says you are finished now.

Children have been reading the second message correctly for fifty years, and the research has spent that whole time catching up. When you want to check whether a child is actually getting better rather than just faster, the trap is the same one described in the illusion of knowing.

Open paperback on a sunlit windowsill with drifting curtain.

Frequently Asked Questions

What is self-determination theory in simple terms?

Self-determination theory is a framework developed by Edward Deci and Richard Ryan proposing that people sustain motivation when three psychological needs are met: autonomy, competence and relatedness. Autonomy is standing behind what you are doing rather than feeling pushed into it. Competence is the experience of being effective at it. Relatedness is feeling connected to people who care about it too. The theory was named in a 1985 book and has since been applied to schools, workplaces, sport and health. Its best known prediction is that conditions which thwart these needs shift behaviour toward external control and motivation drops when the external pressure stops.

Does rewarding a child for reading really make them read less?

It depends entirely on the reward. A 1999 meta-analysis of 128 studies found that tangible rewards promised in advance and given for doing a task reduced voluntary persistence at effect sizes between d = -0.28 and d = -0.40, and that the effect was significantly larger for children than for college students. The same analysis found that verbal praise increased persistence at d = 0.33 and that unexpected rewards had no effect at all. A 2008 experiment with 75 third-graders found that children rewarded with a book scored the same as children rewarded with nothing, while children rewarded with a token scored worse. So a token or a treat carries real risk. A book, a library trip or an honest comment about what improved does not appear to.

What is the overjustification effect?

The overjustification effect is the name for what happens when an external reward is added to an activity someone already enjoys. The person begins to explain their own behaviour by the reward rather than by the interest, and when the reward is removed the reason for the behaviour goes with it. It was demonstrated in 1973 by Mark Lepper, David Greene and Richard Nisbett, who promised 51 preschoolers a certificate for drawing and found that those children later spent about half as much free-play time drawing as children who received the same certificate as a surprise or received nothing. Self-determination theory explains the same result through the controlling aspect of a reward moving the felt cause of the action outside the person.

Are reading programmes like Accelerated Reader or Book It bad for children?

The evidence is genuinely mixed and it is worth separating the outcomes. A randomised trial of Accelerated Reader with 978 students in grades 3 to 6 found reading achievement growing faster in programme classrooms, with effect sizes from +0.07 to +0.34. A 1999 survey of about 107 college students found that having taken part in Book It as a child was associated with neither more nor less reading in adulthood and no difference in intrinsic motivation. Neither study measured what the other one did. The reasonable reading is that these programmes can raise measured achievement without necessarily building or destroying a reader, and that the reward's relationship to reading matters more than the programme's name.

Is praise a reward, and does it undermine motivation too?

Praise behaves differently from tangible rewards in every meta-analysis that has separated them. In the 1999 analysis, positive feedback increased free-choice persistence at d = 0.33 and self-reported interest at d = 0.31. Even Cameron and Pierce's 1994 analysis, which was broadly sceptical of the undermining effect, found verbal praise increasing intrinsic motivation. The important qualification is that the effect for children alone was d = 0.11 and did not reach significance, so praise helps children less than it helps adults. Praise that is specific and informational, describing what actually improved, is closer to competence feedback than to payment, which is why the theory predicts it should not do the same damage.