Introduction
You know this experiment. A small child sits at a table. On the table is one marshmallow. An adult explains the deal, leaves the room, and the rest is famous: wait fifteen minutes and you get two, eat it now and you get one. Then the twist that made it a parable. The children who waited grew up to score higher on the SAT, weigh less, cope better, and generally do better at being adults.
It is the most repeated study in popular psychology. It is also, in almost every version you have read, wrong in a specific and interesting way.
Not wrong because the experiment was faked. It was not. Not wrong because nobody could replicate it either. Two large replications exist, and between them they found a great deal less than the story promises. The problem is narrower and stranger than a fraud story. The marshmallow test is a good measurement. It is just not a measurement of the thing it became famous for.
Here is what happens when you stop reading summaries and open the papers. The founding experiments were not about willpower at all. They were attention experiments, and the researchers were manipulating where the child was looking and what the child was thinking about [1]. The famous correlation with SAT scores comes from one paper in 1990, it holds in one of the study's four experimental conditions, and there were thirty-five children in that condition [2]. In the other three conditions the correlations were negative. That is printed in the original paper, in a table, in plain sight.
And then there is the result that reframes everything. In 2013, three researchers ran the standard test on twenty-eight children, but first they had an adult either keep two small promises or break them. The 14 children who had just watched an adult break a promise waited an average of 3 minutes. The 14 children who had watched the same adult keep one waited 12 minutes [3]. Nothing about the child changed. Five minutes of adult behaviour changed the answer by a factor of four.
This article walks the whole arc, from 1970 to a meta-analysis posted in 2026. Where the popular version is right, it says so. Where it has been flattened into something more satisfying and less true, it takes it apart. And where the field is still arguing, which it is in at least four places, it names who is on which side instead of pretending the matter is closed.
One thing this article will not say is that the marshmallow test was debunked. It was not. Two competent teams have now analysed the same data set and published opposite conclusions, and both are still in print.
What Mischel Was Actually Manipulating
Start at the beginning, because the beginning is not what you think.
Walter Mischel did not set out to measure character. He set out to find out what makes waiting easier, which is a completely different question. The 1970 paper he published with Ebbe Ebbesen is called "Attention in delay of gratification", and the title is the finding [1]. The experiment varied whether the rewards were sitting in front of the child or hidden under a cover. Children who could see the treats waited less time, not more. Attention to the thing you are resisting is the thing that breaks you.
The 1972 follow-up with Ebbesen and Antonette Raskoff Zeiss went further and manipulated what the child was told to think about [4]. Think about how sweet and chewy the marshmallow is, and the wait collapses. Think about something else entirely, or think about the marshmallow as a puffy white cloud, and the wait extends. The child is the same child. The instruction is doing the work.
Read that again, because the popular version inverts it. In the story that circulates, the strong-willed children resist temptation through sheer force. In the actual experiments, the children who lasted were the ones who stopped looking at the marshmallow and found something else to do with their heads. Waiting was not a contest of strength. It was a problem of where to point your attention, and it was solvable with a hint.
Philip Peake, Michelle Hebl and Mischel returned to this in 2002 with a paper on strategic attention deployment, separating what happens when a child is waiting from what happens when a child is working [5]. The strategies that help are not the same in both situations. Distraction helps you wait. It does not help you work.
By 1999, Janet Metcalfe and Mischel had built this into a general account they called the hot and cool systems [6]. The hot system is fast, emotional, reflexive, and triggered by the sight and smell of the thing you want. The cool system is slow, cognitive, and capable of turning a marshmallow into an abstraction. Self-control, in this account, is not the cool system overpowering the hot one. It is the cool system finding a way to stop the hot one from being triggered in the first place.
That distinction matters for everything that follows. If self-control is a muscle, it belongs to the person and travels with them. If self-control is a strategy for managing what your attention lands on, it belongs partly to the situation, and the situation can be changed. The founding papers say the second thing. Almost every retelling says the first.
The mechanism story sits alongside what is known about the prefrontal cortex and its role in holding a goal in mind while something more immediate is shouting for attention. That is real. It is also only part of the picture, and the rest of this article is about the rest of the picture.
The Number That Built the Legend
Now the famous part, and this is where you should slow down.
The claim that the marshmallow test predicts life outcomes comes almost entirely from one paper: Yuichi Shoda, Mischel and Peake, published in Developmental Psychology in 1990 and drawing on 653 children [2]. Those 653 children took part in the original Stanford experiments between 1968 and 1974. Years later the researchers wrote to the parents. Ninety-four of them supplied their child's SAT scores.
Here is the sentence that never survives the retelling. Those 653 children had not all done the same experiment. They had been spread across different conditions, because the original studies were manipulating things. Some waited with the rewards visible. Some waited with the rewards hidden. Some were given a strategy to use. Some were left to work it out themselves.
When Shoda and colleagues looked at the correlation between preschool wait time and adolescent SAT scores, they had to look at each condition separately. Table 4 of the paper prints all four.
Before you read the row that made the test famous, look at how few children are in each of the other three. Then look at the signs.
One cell works. In the condition where the rewards were sitting in plain sight and the child had to invent their own way of coping, wait time showed correlations of 0.42 with SAT Verbal and 0.57 with SAT Quantitative across 35 children [2]. The verbal correlation reached the ordinary threshold for significance. The quantitative one was stronger.
In the other three conditions every single correlation is negative. Not weaker. Negative. In the condition where a strategy was suggested and the rewards were visible, the correlation with SAT Verbal was minus .40, on fourteen children.
The authors are honest about this. They write that in some conditions the sample sizes "became barely sufficient for a meaningful computation of correlations". They are not hiding anything. The information is right there, and it has simply never made it into the retelling, because the retelling wants one number and the paper has sixteen.
There is more in the footnotes. Inside that one working condition, the SAT Verbal correlation was r = .02 across 17 boys and r = .74 across 18 girls, and the authors tested the difference and found it significant [2]. Taken at face value, that would mean the entire verbal finding comes from the girls. Taken sensibly, it means seventeen is not enough children to say anything about boys. A 2018 meta-analysis of twenty-eight papers and fifty-two effect sizes found the sex difference in delay of gratification among typically developing samples to be negligible [7]. The 1990 split is almost certainly noise. It is in this article only so that you know it exists and know why it goes nowhere.
None of this makes the 1990 paper bad. It was a small, careful longitudinal follow-up that reported its own limits. What happened afterwards was not the paper's fault. A finding based on thirty-five children in one experimental condition escaped into the world as a general law about human character, and stayed there for twenty-eight years.
What Happened When Somebody Ran It Again
In 2018, Tyler Watts, Greg Duncan and Haonan Quan did the obvious thing that nobody had done. They ran it again, on 918 children who were a great deal less unusual than the Stanford sample [8].
Their data came from the National Institute of Child Health and Human Development Study of Early Child Care and Youth Development. The analytic sample was 918 children who had a valid delay-of-gratification measure at fifty-four months and complete achievement and behavioural data at fifteen. The headline analysis focuses on the 552 children whose mothers had not completed college, a group the paper points out is about ten times the size of the 1990 sample.
The correlation was there. It was also about half the size of the original, and it lost roughly two thirds of what remained once the analysis controlled for family background, early cognitive ability and the home environment [8]. In the version that survived those controls, one additional minute of waiting at age four predicted about a tenth of a standard deviation of achievement at age fifteen. Associations with behavioural outcomes at fifteen were much smaller and rarely statistically significant.
Then there is the detail from that paper that deserves more attention than it gets. Most of the variation in adolescent achievement among those 918 children came from being able to wait at least 20 seconds [8].
Twenty seconds. Not fifteen minutes. If nearly all of the predictive signal is in the first twenty seconds, then whatever the task is picking up, it is not a capacity for sustained endurance. It is something that shows up almost immediately and then stops adding information. A child who can hold still for twenty seconds and a child who holds still for twelve minutes are not, on this evidence, in meaningfully different categories.
That chart is from a different study, on 28 children, and we will come to it shortly. It is here for one reason. The 918-child replication had to work hard to move achievement by a tenth of a standard deviation per minute of waiting. Five minutes of adult behaviour moved the waiting itself by nine minutes. Those two numbers are not on the same scale and cannot be compared directly, but they tell you where the large effects in this literature live.
And Again, All the Way to Twenty-Six
Fifteen is a school-leaving age. It is not an outcome, and it is not where the original claim pointed.
In 2024, Jessica Sperber, Deborah Lowe Vandell, Duncan and Watts published the answer in Child Development, following 702 participants to the age of twenty-six [9]. Those 702 people, 83 percent White and 46 percent male, took the marshmallow test at fifty-four months in 1995 and 1996, and completed survey measures at age twenty-six in 2017 and 2018. The analysis was preregistered, which means the researchers committed to what they would test before they looked.
The result is close to nothing. Across the 702 participants there were correlations of 0.17 with educational attainment and minus 0.17 with body mass index, and almost every regression-adjusted coefficient was not statistically significant [9]. There was no clear pattern of the effect being stronger for poorer children, or for boys, or for girls.
A correlation of .17 is not zero. It is also not a life sentence. It means that if you lined up every child by how long they waited and lined them up again by how far they got in education, the two lines would agree a little more often than chance. It is the kind of association that is real at the level of a population and useless at the level of a person.
That last point is worth holding onto, because it is where most of the damage in this story was done. Even in the most generous reading of the original findings, the marshmallow test never told you anything reliable about an individual child. It described a weak tendency across a group. Every version that treats it as a diagnosis for one four-year-old was already overreaching before the replications arrived.
Compare the three studies directly and the shape of the field becomes clear.
The Same Data, Two Conclusions
This is the part where a tidy story would end, and it does not end here.
In 2020, Laura Michaelson and Yuko Munakata published a paper in Psychological Science with a title that says everything: "Same Data Set, Different Conclusions" [10]. They ran an independent, preregistered secondary analysis of the same data the 2018 replication used, applying the analytic approach of the original marshmallow studies rather than the newer one. They found significant associations for three of the five outcomes they tested. Relationships between delay and problem behaviour held in bivariate, multivariate and multilevel models, where the other analysis had found none.
Two competent teams. One data set. Opposite conclusions. This is not a scandal, it is what a live methodological disagreement actually looks like, and it is why the word "debunked" does not belong in this story.
Michaelson and Munakata also propose an interpretation, and it is the interpretation this whole article has been circling. The relationships they found were better explained by social support than by self-control. Their reading is that the marshmallow test predicts things because it reflects aspects of a child's early environment that matter over the long term [10]. The test is picking up a signal. The signal is coming from outside the child.
Watts and Duncan answered their critics directly in 2019, in the same journal, in a paper about controlling, confounding and construct clarity [11]. You can read both sides as the participants themselves wrote them, which is rarer in psychology than it should be.
So where does that leave you? With a small effect, fragile to analytic choices, and mostly explained by the circumstances the child arrived with. That is a less exciting sentence than "the marshmallow test predicts success". It is the one the evidence supports.
Five Minutes That Change Everything
Now the experiment that should be as famous as the original and is not. If you remember one thing from this article, make it this one.
In 2013, Celeste Kidd, Holly Palmeri and Richard Aslin published a paper in Cognition, with the title "Rational snacking", built on 28 children [3]. They were aged between three and a half and nearly six, fourteen in each condition, balanced for gender and age. It is a small study and this article will keep saying so.
Before the marshmallow test, every child did an art project with an adult. The adult made two small promises. In one condition the adult kept both: they said they would come back with better art supplies, and they did. In the other condition the adult broke both: they came back apologetic and empty-handed, twice.
Then the marshmallow test ran exactly as normal.
The 14 children who had just been let down twice waited an average of 182 seconds, which is three minutes and two seconds. The 14 who had just been kept faith with waited an average of 722 seconds, which is twelve minutes and two seconds [3]. Only 1 out of 14 in the unreliable group made it to the full fifteen minutes, against 9 out of 14 in the reliable group [3]. The difference held on both the timing and the count.
Think about what that does to the interpretation. Nothing about the children differed. They were randomly assigned. The only variable was five minutes of an adult's behaviour immediately beforehand, and it produced a four-fold difference in a measure that had been read for forty years as a stable property of the child.
The children in the unreliable condition were not weak. They were correct. They had just accumulated direct evidence that promises from this adult in this room do not pay, and they acted on it. Eating the marshmallow is the rational move when the second marshmallow is unlikely to arrive.
This is not a one-off. Michaelson and colleagues showed in 2013 that delaying gratification depends on social trust [12], and Michaelson and Munakata followed it in 2016 with a result that is somehow even sharper: preschoolers wait less after simply watching an adult treat a third person badly [13]. The child does not need to be the victim. Observing unfairness is enough. A 2022 study replicated the trust effect in Costa Rican preschoolers [14].
There is a version of this that a reader can feel from the inside. You have worked somewhere where the promised bonus never quite materialised. After the second time, you stopped planning around it. Nobody would call that a character defect. It is updating on evidence, and it is the same computation a four-year-old is running at the table.
The diagram makes the substitution visible. What gets written down at the end is a number of seconds, labelled self-control. What produced the number includes an estimate the child made about somebody else.
The Test Travels Badly
If the marshmallow test measured a general human capacity, it would behave the same way wherever you took it. It does not.
In 2018, Bettina Lamm, Heidi Keller, Johanna Teiser, Helene Gudi and colleagues published the first study to run the task on children outside the Western middle class, comparing 125 German children with 76 children from the rural Nso farming community in Cameroon [15]. All of them were four years old. The treat was adjusted to be locally desirable in each place, which is the obvious thing to do and had somehow not been done before.
The Nso children outperformed the German children.
If your instinct is to look for what went wrong with the German sample, notice that instinct. It is the assumption the whole test was built on, that the Western middle-class result is the baseline and everything else is a deviation from it.
That result has a way of getting told as a curiosity, and it is not a curiosity. It is a direct challenge to the idea that the task indexes a general developmental capacity, because the group with less formal schooling, fewer material resources and no exposure to the psychological tradition the test came out of did better at the test.
The second study in the same paper asked why, and the answer is about parenting rather than children. Nso mothers' emphasis on hierarchical relational socialisation goals and responsive control was associated with better delay performance than German middle-class mothers' emphasis on psychological autonomy and sensitive child-centred parenting [15]. Different cultures teach different things about waiting, and the children learn what they are taught.
You may have seen specific percentages attached to this study. Numbers in the region of 70 percent of Cameroonian children waiting against 28 percent of German children circulate widely. Those figures come from press coverage rather than from the paper's own abstract, and this article does not print numbers it cannot trace to the source. The direction of the result is solid and published. The precise percentages are somebody's summary, and summaries drift.
A separate comparison in 2021 by Ning Ding, Anna Frohnwieser, Rachael Miller and Nicola Clayton tested 75 children in China aged three to five against a previously published sample of 61 children in Britain on a delay-choice task, and the children in China performed better once reward visibility was manipulated [16]. Different task, same direction of surprise.
Food, or Gifts
The cross-cultural work gets sharper still, and the next study is the one that should change how you describe the task.
In 2022, Kaichi Yanaoka, Michaelson, Ryan Mori Guild, Grace Dostart and colleagues ran a preregistered study in Psychological Science on 80 children in Japan and 58 children in the United States [17]. Each child did the delay task twice: once waiting for food, once waiting for a wrapped gift.
Japanese children waited longer for food than for gifts. American children waited longer for gifts than for food.
Ask yourself which one you would find easier. Whatever your answer is, it probably tells you where you grew up rather than how much self-command you have.
The same children. The same task. Opposite results, and the direction tracks exactly what each culture practises. Japanese children routinely wait before eating, with a set phrase said at the start of a meal. American children routinely wait to open presents, at birthdays and at Christmas. Whatever the child has rehearsed is the thing they are good at waiting for.
That finding does something quite specific to the willpower interpretation. A general capacity does not switch on for marshmallows and off for wrapped boxes depending on which country you are in. A habit does.
Yanaoka, Michaelson, Satoru Saito and Munakata pushed on this in 2026 with 149 children in Japan aged four to six, and found the part that makes the habit account convincing [18]. Children with stronger habits of waiting to eat waited longer for food, as expected, but they also reported that waiting took less work. Not more resistance. Less resistance needed.
Read that slowly. A habit does not make you better at fighting an urge. It removes the urge you would otherwise have had to fight. The child who has said the phrase before every meal for three years is not exercising restraint at the table. They are doing the normal thing.
In the same study, the experimenter's trustworthiness did not significantly change how long children waited, which is worth reporting because it does not match the Kidd result. What it did change was how much work the children said the waiting took, with the untrustworthy experimenter making the wait feel harder. Yanaoka and colleagues have also written a review setting out the case for effortless control as the mechanism [19]. A 2023 study from the same group found that a preschooler's delay is also shifted by what children in their own group are seen doing [20].
This connects the marshmallow literature to something you can read about separately in how habits become automatic. A rehearsed behaviour stops costing what it used to cost. The saving is real and it is invisible from outside, which is exactly why the child who has it looks like they have more willpower.
Waiting Is a Bet, Not a Virtue
Everything so far points the same way, and it is worth stating the underlying logic directly.
When a four-year-old decides whether to wait, they are not consulting their reserves of moral fibre. They are making a forecast. Will the second marshmallow arrive? How long is this going to take? What are the odds this adult is telling the truth? The wait is a bet, and the size of the bet depends on the odds.
Joseph McGuire and Joseph Kable made this formal in 2013, showing that rational temporal predictions can underlie what look like failures to delay [21]. If you do not know when the reward is coming, and the time you have already waited tells you the wait might be very long, then quitting is the correct policy under many plausible assumptions. Impatience, in that framing, is not a defect of the person. It is what an optimal agent does when it has bad information about timing.
That reframing explains a whole cluster of findings that otherwise look like a catalogue of deficits.
Vladas Griskevicius, Joshua Tybur, Andrew Delton and Theresa Robertson showed in 2011 that cues of mortality shift preferences toward immediate rewards, and that the shift depends on childhood socioeconomic status [22]. In their life-history framing, discounting the future steeply is not an error when the future is genuinely unreliable. Lei Liu, Tingyong Feng, Tao Suo and Kang Lee found in 2012 that simply exposing adults to poverty cues pushed them toward short-term choices [23].
More recent work has taken the same idea apart into its components. A 2023 study by Liang Yu and colleagues traced the same effect from perceived scarcity through self-efficacy [24], and a 2026 experiment using a serious game found that uncertainty amplifies what scarcity already does [25]. Gary Evans and Jeyon Kim reported in 2024 that the relation between substandard housing and children's self-regulation is worse for some children than others [26].
The pattern in all of them is the same. Put a person in a world where the future is less certain, and they will discount the future more. That is not a personality trait leaking out. That is a reasonable response to a real situation, measured in a laboratory that mistakes it for a trait.
Which raises the uncomfortable question about the original longitudinal findings. If poorer children wait less because waiting pays off less reliably in their lives, and poorer children also have worse educational outcomes for reasons that have nothing to do with marshmallows, then a correlation between wait time and later achievement is exactly what you would expect even if the wait time itself caused nothing at all. That is precisely the confound the 2018 replication was built to test, and it is why the effect shrank by two thirds when the controls went in.
When the Waiting Is Somebody Else's Problem Too
There is a version of the task that almost nobody runs, and it produces the largest improvement in the literature.
In 2020, Rebecca Koomen, Sebastian Grueneisen and Esther Herrmann published a study in Psychological Science that paired 207 children up and made the reward interdependent [27]. They were drawn from two very different places: Germany and Kenya. In the cooperative version, both children only got the larger reward if both of them waited. If either one gave in, neither got it.
Children performed substantially better on the cooperative version than on the standard one, in both countries [27].
The child who cannot wait for themselves can wait for a partner. That is a strange result if the task measures a fixed capacity, and an unsurprising one if it measures motivation shaped by the social situation. It also suggests something about how the original test was framed, which is as a private negotiation between a child and their own appetite. Very little in a child's actual life is structured that way.
Koomen returned to the question in 2025 with a study run entirely online during and after the pandemic, testing on 66 children whether an explicit promise from a partner changes anything [28]. They were aged five and six, in the United Kingdom, taking part from their own homes over video call, interacting with a confederate child who either promised not to eat his treat or raised the possibility that he might. Children in the promise condition waited longer.

Koomen R, Waddington O, Goncalves LSM, Köymen B, Jensen K. Does promising facilitate children's delay of gratification in interdependent contexts?. R Soc Open Sci.; 12(5):250392. https://doi.org/10.1098/rsos.250392. Figure 2.. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.
Look at the spread in that figure before you look at the difference between the two conditions. Both distributions run from almost zero to the full six hundred seconds. Some children in the social risk condition waited the entire time. Some children in the promise condition ate immediately. The condition shifts the average, and the average sits inside an enormous range of individual variation that no condition explains.
That picture is the honest summary of this entire literature, and it is worth carrying to every other claim in this article. The experimental effects are real. They are also modest against how differently individual children behave in the same room under the same instruction.
What the Whole Literature Says at Once
One study can always be an accident. The way you find out is by counting all of them.
Individual studies are one thing. In 2026, Mehmood Ul Hassan Ul Hassan Shajih and Sabine Doebel gathered the field into a single meta-analysis, synthesising 119 samples from 86 published and unpublished studies with maximum waiting periods ranging from seven to twenty-five minutes [29]. This one is a preprint and has not yet been through peer review, so treat it as strong evidence rather than settled evidence.
Several of its findings are the ones you would guess. Older children waited longer and were more likely to complete the full delay. Samples characterised as lower socioeconomic status had shorter mean wait times.
Several of them are not what you would guess at all.
Socioeconomic status predicted shorter average waits but did not predict whether a child made it to the end of the full delay period [29]. Allowing children to choose their own reward made no difference to waiting, and neither did explaining to the child why the experimenter was going away, with moderate Bayesian evidence for the null on both. Most other procedural details the researchers coded were also unrelated to waiting.
The result with the most practical bite is about the length of the test itself. Longer maximum durations produced longer average waits, but lower odds of any given child completing the full delay, and higher rates of children being excluded from the analysis [29]. Three different laboratories using seven, fifteen and twenty-five minute versions are not running the same experiment, and the differences between their results may be partly an artefact of the clock.
One more finding from that meta-analysis leads directly into the next section. Publication year was positively associated with mean wait time, in several models. Children in more recent studies waited longer.
Children Are Getting Better at This
Here is the finding that should have made more noise than it did.
Before running his analysis, John Protzko polled 260 experts in cognitive development and asked what they expected. Fully 84 percent of them believed that children today are worse at delaying gratification than children of the past, or at best no different [30]. That is a near consensus, and it matches what you have almost certainly absorbed from a hundred articles about attention spans and short video.
Then he analysed fifty years of marshmallow test data.
Delay of gratification times have been increasing, at roughly a fifth of a standard deviation per decade [30]. Protzko points out that this mirrors the magnitude of the secular gains in IQ scores recorded over the same kind of timescale.
Children in 2020 waited longer than children in 1970. Not shorter. Longer. Through the arrival of television, video games, the internet, the smartphone and the algorithmic feed.
This does not prove that screens are harmless, and it is not offered as proof of that. It proves something narrower and more useful: the specific claim that modern life has eroded children's capacity to wait, which is stated as obvious in a great many places, is contradicted by the only long-run measurement anyone has made of it. Not one of those 84 percent of specialists expected an increase. That is worth remembering the next time a claim about kids these days feels self-evident.
The Brain Result, and Its Twenty-Six People
At some point in almost every retelling, brain imaging arrives to settle the argument. It does not settle it, and the reason is a number.
In 2011, B. J. Casey, Leah Somerville, Ian Gotlib, Ozlem Ayduk and colleagues tracked down nearly sixty adults from the original Stanford cohort, by then in their mid-forties, and tested them on hot and cool versions of a go/nogo task [31]. The people who had been low delayers as preschoolers, and who had also shown low self-control in their twenties and thirties, performed worse when they had to suppress a response to a happy face. Not to a neutral face. Not to a fearful one. A happy one.
Twenty-six participants from that group went into a scanner [31]. In that subset, the prefrontal cortex distinguished between go and nogo trials more sharply in the high delayers, while the ventral striatum was more strongly recruited in the low delayers.
Twenty-six people. That is the imaging sample behind forty years of headlines about self-control being visible in the brain.
You have probably seen one of those headlines. You have almost certainly never seen that number attached to it.
The result is interesting and it is probably pointing at something real. It is not a foundation. An imaging study on twenty-six participants is a lead worth following, and the field has followed it. Michelle Achterberg and colleagues showed in 2016 that frontostriatal white matter integrity predicts the development of delay of gratification over time, in a longitudinal design that is better suited to the question [32]. Margaret Benningfield and colleagues found in 2014 that caudate responses to reward anticipation track delay discounting in healthy young people [33].
The imaging field has since grown large enough to be summarised. A 2026 meta-analysis pulled together the neural systems underlying delay discounting across the imaging literature [34], and a 2023 review covered what brain imaging and non-invasive stimulation have found together [35].
That is the honest position on the neuroscience, and it is less dramatic than the headlines. There is machinery involved in waiting, the machinery differs between people, and nobody has shown that the differences are fixed.
There is also causal evidence now, which correlational imaging can never provide. A 2025 study using frontoparietal theta stimulation linked working memory to impulsive choice directly, by intervening rather than observing [36]. The dopamine side of the same decision has its own literature, and a 2025 review sets out its role in weighing risk against reward [37]. If you want the fuller version of that mechanism, it is covered in the article on dopamine and learning.
None of this contradicts anything earlier in this piece. A child's forecast about whether the second marshmallow is coming has to be computed somewhere, and the frontostriatal circuitry is a reasonable candidate for where. Finding the machinery does not tell you what the machinery is calculating.
Two Words That Are Not the Same Thing
A confusion runs through almost every popular account of this research, and through a fair amount of the research itself. Two different measures get treated as one, and separating them takes a paragraph.
Delay of gratification is the marshmallow situation. A reward is in front of you, you have already been given it, and the question is whether you can leave it alone for a period of time. It is a test of maintaining a decision under continuous temptation.
Delay discounting is a choice, usually made once, on paper or a screen. Ten pounds now or fifteen pounds in a month. Nothing is sitting in front of you. Nothing has to be endured. The question is how steeply you devalue a reward as it moves further into the future.
Elsa Addessi, Fabio Paglieri and Michael Beran demonstrated in 2013 that these come apart empirically, distinguishing delay choice from delay maintenance in capuchin monkeys and showing the two measures do not rank individuals the same way [38]. Brady Reynolds and Ryan Schiffbauer had proposed a unifying feedback model in 2005 that treats them as two views of one process [39]. That disagreement is not resolved.
Lars Göllner and colleagues examined in 2018 how both relate to age, to episodic future thinking and to how far ahead a person habitually looks [40]. Their answer is that the two measures share some ground and not all of it.
Those two constructs are usually described as measuring the same thing at different ages. They do not.
A fair amount of recent work in this corner is about making the field comparable to itself. Sara Garofalo and colleagues published normative delay-discounting data with an open task analysis tutorial in 2022, which is the kind of unglamorous methodological work that makes a field comparable across laboratories [41]. June Pilcher, Drew Morris and Dylan Erikson reviewed self-control measurement methods across the board in 2022 and found the same construct being measured in ways that do not agree with each other [42].
The practical consequence for you as a reader is simple. When a headline says a study found that self-control predicts something, check which task they used. A monetary discounting questionnaire in adults and a marshmallow in front of a four-year-old are not interchangeable, and results from one do not transfer cleanly to the other.
Is It Even Self-Control?
The construct question goes deeper than measurement, and Angela Duckworth put it directly in a 2013 paper titled "Is It Really Self-Control?" [43].
Duckworth, Eli Tsukayama and Teri Kirby ran two studies. The first used 56 children of school age and found that delay time was associated with teacher ratings of self-control and with Big Five conscientiousness, but not with intelligence and not with reward-related impulses [43]. The second used 966 children of preschool age and found delay time consistently associated with parent and caregiver ratings of self-control, again not with reward-related impulses [43].
Their conclusion is careful. Delay task performance can be influenced by traits that have nothing to do with self-control, but its predictive power comes primarily from the self-control it does capture [43]. What predicted later academic, health and social outcomes most consistently, though, was not the delay time itself. It was the ratings of effortful control that adults gave the same children.
That is a quiet but important result. If a teacher's judgement of a child predicts the child's future better than a fifteen-minute laboratory task does, then the laboratory task is at best a noisy proxy for something a person can see across a whole school year.
The older literature had already got close to this from the other direction. David Funder and Jack Block reported in 1989 that ego-control, ego-resiliency and IQ each contributed to delay of gratification in adolescence [44], and an earlier paper with Jeanne Block in 1983 had traced longitudinal personality correlates of delay [45]. Robert Krueger, Avshalom Caspi, Terrie Moffitt and Jennifer White asked in 1996 whether low self-control is specific to any one form of psychopathology and found it is not [46].
The picture that emerges is of a task that touches several things at once: attention, temperament, trust, habit and a forecast about the world. That is not a flaw in the task. It is a flaw in the label attached to it.
What Actually Moves It
If the wait is partly a strategy, then teaching the strategy should move the number. It does.
This is the part of the literature that is genuinely useful, and it is also the part that gets least attention, because it is undramatic. Nobody writes a bestseller about telling a child what to think about.
Caterina Gawrilow, Peter Gollwitzer and Gabriele Oettingen showed in 2010 that if-then plans improve delay of gratification performance in children both with and without ADHD [47]. The child decides in advance what they will do when the urge arrives, and the decision does the work in the moment rather than the willpower.
Joanne Murray, Anna Theakston and Adrian Wells asked in 2016 whether the attention training technique could turn one marshmallow into two, which is one of the better paper titles in this field, and found that training attention improved children's ability to wait [48]. Daniel Romer, Duckworth, Sharon Sznitman and Sunhee Park had reported in 2010 that adolescents can learn self-control, in the context of controlling risk taking [49].
Both of those results have the same shape. The child is not being made stronger. The child is being handed a plan.
A separate line of work uses the future rather than the present. Patrick Burns, Cristina Atance, Patrick O'Connor and Teresa McCormack cued children and adolescents to imagine specific future events and found it reduced delay discounting [50]. Ciarán Canning, Agnieszka Graham and McCormack extended reward-related episodic future thinking to delayed gratification in children in 2023 [51], and a 2025 individual-participant analysis pooled the evidence across those studies [52].
The method keeps working when it is made simpler. Katelyn Carr and Leonard Epstein reported in 2026 that written or drawn cues for episodic future thinking improve discounting in children [53]. Burns and colleagues had already established in 2021 that thinking about the future and delaying gratification travel together in children [54].
The next two go further than a single trick.
You can go one level up from teaching a specific trick and teach the habit of looking for one. Patricia Chen and colleagues found in 2026 that a strategic mindset, meaning the habit of asking what strategy might work here, improves children's generation of effective strategies [55]. Cansu Tutkun had reported in 2022 that the distraction strategies preschoolers already use predict how they behave more broadly [56].
The cheapest intervention in the whole literature costs one sentence. Rachel Karniol and colleagues showed in 2011 that asking a child to imagine being a character known for patience changes how long they wait, in a paper titled "Why Superman Can Wait" [57]. The child is not stronger for being Superman. They just have somewhere else to put their attention.
Notice what all of these have in common. Every intervention that works is a strategy intervention. Nothing on that list strengthens a faculty. They all give the child something specific to do instead, which is exactly what Mischel was manipulating in 1970 with his instruction to think about a puffy white cloud. Fifty-six years of research on how to improve delay of gratification has closed a loop back to the founding experiment.
There is an obvious implication about the direction of causation that follows from this and it is worth naming. If wait time can be moved by an instruction given thirty seconds earlier, then wait time measured on one afternoon is not a stable trait. It is a snapshot of what strategy the child happened to have available that day.
The same logic runs through what is known about procrastination as mood repair, where the failure to start something is better explained by what the person is trying to avoid feeling than by any deficit of discipline.
Parrots, Crows and a Bumble Bee
A short section, because it constrains the story more than its length suggests.
Waiting for a better reward is not a human speciality. Michael Beran documented self-imposed delay of gratification in four chimpanzees and an orangutan in 2002 [58], and Beran and William Hopkins reported in 2018 that chimpanzee self-control relates to general intelligence [59]. Alexandra Schnell, Markus Boeckle and Clayton reviewed delay of gratification in corvids in 2022 and its relationship to other cognitive abilities [60].
Birds are the surprise in this literature. Irene Pepperberg and Virginia Rosenberger showed in 2022 that a grey parrot will wait for more tokens [61]. Rachael Miller and colleagues compared crows, parrots and nonhuman primates in a 2019 review [62].
The furthest edge of this literature is a 2025 paper testing the limits of delay of gratification in bumble bees [63]. Robin Dunbar and Susanne Shultz argued in the same year that self-control has a social role in primates that it does not appear to have in other mammals or birds [64], which is a more interesting claim than a simple ranking of species by patience.
The point here is a corrective one. Whatever a marshmallow test measures, it is not a uniquely human moral capacity that some children have more of. A capacity that turns up, in some form and within limits, in a bumble bee is not a moral quality.
Where the Test Lives Now
A finding does not stop being used because somebody questions it. It moves.
The marshmallow task did not go away when the replications landed. It moved into clinical and applied research, and that is where most current work using it sits.
The clinical use is the largest of them, and it is the one where the measurement has the clearest job. Connor Patros and colleagues published a meta-analysis in 2016 on choice-impulsivity in children and adolescents with ADHD, finding a consistent preference for smaller immediate rewards [65]. Girija Kadlaskar and colleagues looked at delay of gratification in preschoolers with autism and concerns for ADHD in 2024 [66].
Notice that none of this rests on the task predicting anything decades later. It rests on the task describing a difference now, which is a much more modest and much better supported use of it. Jeffrey Gagne's 2017 synthesis covers how self-control develops in early childhood across these literatures [67].
Screens are the other place the task keeps turning up, and the evidence there is much weaker than the coverage suggests.
There is a large body of correlational work on the subject. Henry Wilmer and Jason Chein found in 2016 that patterns of mobile device use are associated with intertemporal preferences [68], and a 2023 study linked online gaming and short video behaviour to academic delay of gratification [69].
Handle that last group carefully. These are correlations measured at one point in time, and they sit next to Protzko's fifty-year analysis showing that children's delay times have been rising throughout the period when all of those technologies arrived. Both can be true. Heavy users of short video may well differ from light users in ways that show up on a discounting task, without short video having made the population as a whole less patient. Cross-sectional differences and long-run trends are different questions, and conflating them is how a moral panic gets a citation.
For what it is worth, the relationship between waiting, reward and compulsive behaviour has a serious literature of its own, covered in more depth in the piece on how addiction hijacks the learning system.
Parenting research has kept using the task too. Teresa Jacobsen and colleagues connected children's ability to delay to mother-child attachment as far back as 1997 [70]. Pauline Effenberger and colleagues examined the relationship between parental personality and children's delay ability in 2021 [71].
The most recent work has followed the consequences further out. Heather Leonard, Atika Khurana and Derek Kosty traced early parenting effects on childhood delay and adolescent physiological stress load in 2025 [72], and Hüseyin Kotaman and colleagues looked at temperament, parenting and teacher self-control together in the same year [73].
None of that parenting work claims to have found a cause. It is all correlational, and the direction could run either way: a calmer child may produce calmer parenting rather than the reverse.
One question the field is still working on is how much of any of this holds still as a child grows. Ariadne Brandt and colleagues published work in 2025 on the stability and change of cool and hot executive functions across middle childhood [74], and Alexandra Hendry and colleagues examined inhibitory control and problem solving in early childhood in 2022 [75]. If those capacities are still moving at eight and ten, a measurement taken at four was never going to be destiny.
Waiting and staying with something difficult are not the same skill, but they rhyme. The second one is covered separately in the piece on deep focus and why your brain resists it.
The Argument, in Order
Fifty-six years of a single experiment, laid out.
Read it as a story rather than a list. You are watching a claim get built, get tested, and get replaced by a better one.
Read down that column and the shape of the thing is obvious. The first thirty years build a claim. The next thirteen take it apart and put something better in its place. The last eight are about what the task really responds to, which turns out to be trust, culture, habit and the odds.
What the Test Is Actually For
So what should you take from all of this?
Not that the marshmallow test is worthless. It is one of the most productive experimental paradigms in developmental psychology, and it is still generating results in 2026 that change how the field thinks. A task that can be moved by a broken promise, a wrapped gift, a partner's word and an instruction to imagine a cloud is a task that is sensitive to a great deal. Sensitivity is a virtue in a measurement.
Not that it was debunked either. Michaelson and Munakata found significant associations for problem behaviour in the same data set where the 2018 analysis found none, and theirs was preregistered. The disagreement is about analytic choices, not about fraud, and it has not been resolved.
What you should take is that the label was wrong. The task was called a measure of self-control, and it was written up as though the number belonged to the child. The evidence of the last fifteen years says the number belongs partly to the room. It belongs to whether the adult in front of the child has been reliable, to what the child's culture has taught them to wait for, to how certain their world has been about delivering on promises, and to what strategy happens to be available to them that afternoon.
If you have carried around the idea that you either have willpower or you do not, and that this was settled by an experiment with marshmallows in the 1960s, the honest correction is smaller and more useful than either the myth or the takedown. Nobody has ever demonstrated that waiting is a fixed property of a person. What has been demonstrated, repeatedly, is that waiting is easier when the wait is worth it, when you have done it before, and when you have something to do with your attention while it passes.
That is not a reassuring story about character. It is a fairly practical one about circumstances, and circumstances can be changed. The children who did best were not the strongest. They were the ones who had been given the best reasons to believe.
Frequently Asked Questions
What did the marshmallow test actually measure?
The original experiments measured how attention and strategy affect waiting, not willpower. Walter Mischel and Ebbe Ebbesen showed in 1970 that children waited less when the rewards were visible, and a 1972 study showed that changing what the child was told to think about changed the wait again. Later work has added trust, habit and expectation. A 2013 study by Celeste Kidd and colleagues found that children who had just watched an adult break two promises waited an average of three minutes while children who had watched the same adult keep them waited twelve. The task is best understood as measuring a child's forecast about whether waiting will pay off, combined with whatever strategy they happen to have available.
Has the marshmallow test been debunked?
No, and the word is misleading. A 2018 conceptual replication by Tyler Watts, Greg Duncan and Haonan Quan using 918 children found the effect was about half the original size and shrank by roughly two thirds once family background and early cognitive ability were controlled for. A preregistered 2024 follow-up to age twenty-six with 702 participants found almost no significant adjusted associations. But in 2020 Laura Michaelson and Yuko Munakata analysed the same data set with the original analytic approach and found significant associations for three of five outcomes. Two competent teams reached opposite conclusions from identical data, and the disagreement is still live.
Does delayed gratification in childhood predict success in adulthood?
On the best current evidence, barely. The 2024 preregistered study following 702 people to age twenty-six found raw correlations of .17 with educational attainment and minus .17 with body mass index, and almost all regression-adjusted coefficients were not statistically significant. Even in the original 1990 study the famous SAT correlations held in only one of four experimental conditions, with thirty-five children in it, and were negative in the other three. The association may exist at the level of a population. It has never been shown to be informative about an individual child.
Why do children in some cultures wait longer than others?
Because different cultures practise waiting for different things. In 2018 Bettina Lamm and colleagues compared 125 German middle-class preschoolers with 76 children from the rural Nso farming community in Cameroon and found the Nso children performed better, with the difference linked to maternal socialisation goals. In 2022 a preregistered study found that Japanese children waited longer for food than for a wrapped gift while children in the United States did the opposite, which tracks the local custom of waiting before eating in Japan and waiting to open presents in the United States. A 2026 follow-up found that children with stronger waiting habits also reported the waiting took less work.
What is the difference between delay of gratification and delay discounting?
Delay of gratification is the marshmallow situation. The reward is present, you already have it, and you have to leave it alone under continuous temptation. Delay discounting is usually a single choice made on paper or a screen between a smaller amount now and a larger amount later, with nothing physically present. A 2013 study by Elsa Addessi and colleagues showed the two measures do not rank individuals the same way, distinguishing delay choice from delay maintenance. Results from one do not transfer cleanly to the other, which matters when reading headlines about self-control research.
Can delayed gratification be taught?
Every intervention that reliably works is a strategy intervention rather than a strength intervention. If-then plans improved performance in children with and without ADHD in a 2010 study, attention training improved waiting in a 2016 study, and cueing children to imagine specific future events reduced delay discounting in several studies from 2021 onward. A 2011 study found that asking children to imagine being a character known for patience changed how long they waited. None of these strengthen a faculty. They give the child something else to do with their attention, which is precisely what Mischel was manipulating in 1970.




