The Christmas Experiment
German university students were asked to name a project they meant to finish over the Christmas vacation. Not a resolution. A specific thing, like writing a seminar paper or settling a family argument. Some were easy. Some were the kind people carry around for months without touching.
Half of them were then asked to do one extra thing. They wrote down when and where they would start, in one sentence.
That was the manipulation. No motivation speech, no accountability partner, no reward.
When the researchers followed up, the students who had written the sentence had completed 62 percent of their difficult projects. The ones who had not completed 22 percent [1]. In a second study 17 people wrote down when and where they would write a report about Christmas Eve, and 12 of them did it within two days. Of the 19 who did not write the sentence, 6 managed it. Sixty-eight students in the first study, thirty-six in the second, and a difference big enough that people are still citing it nearly thirty years later.
The technique has a name. Peter Gollwitzer introduced it in 1993 and called it an implementation intention [2]. Everyone else calls it an if-then plan.
You have probably met it without the label. If I get home from work, then I will change into running clothes. If it is Tuesday at seven, then I will open the problem set. The form never varies. A situation on the left, an action on the right, decided in advance and left alone.
This article is about what that sentence does, what it does not do, and why the number attached to it for twenty years is bigger than the evidence now supports. The correction did not come from a critic. It came from the people who built the technique.
Two Kinds of Intention
Start with the distinction, because almost nothing written about if-then plans draws it clearly and the whole thing collapses without it.
A goal intention is a statement about what you want. I intend to get fit. I intend to pass this module. It names an outcome and stops.
An implementation intention is a statement about a moment. If situation Y happens, then I will do X. It names a cue and an action and welds them together.
The second is not a stronger version of the first. It is a different sentence doing a different job, and it sits on top of the goal rather than replacing it. Gollwitzer laid out the argument in the 1999 paper that made the idea famous [3]. The claim was that most of the failure in goal pursuit happens at the moment of starting, not at the moment of wanting, and that you can hand the starting over to the environment.
That is the sentence worth sitting with. You hand the starting over to the environment.
Willpower, in the ordinary sense, is what a moment costs you when it arrives and you still have to decide. The if-then plan removes that moment. If the decision is already made and welded to a cue, then when the cue shows up there is nothing left to decide. It does not beat willpower by giving you more of it. It beats willpower by arranging things so you never reach for it.
The Gap the Plan Is Meant to Close
There is a specific problem this was invented to solve, and it has been measured.
In 2006 Thomas Webb and Paschal Sheeran pooled 47 experimental tests in which researchers had successfully changed people's intentions and then watched what happened to their behaviour. A medium-to-large shift in intention, d = 0.66, produced only a small-to-medium shift in behaviour, d = 0.36 [4]. Roughly half the movement was lost between deciding and doing.
Psychologists call it the intention-behaviour gap. You know it as January.
The early studies went straight at that gap. In 1997 women who specified when and where they would perform a breast self-examination were substantially more likely to have done it a month later, and the plan also weakened the usual link between how strongly someone intended to act and whether they acted [5]. That second finding is the interesting one. If a plan works by making intention matter less, something other than motivation is carrying the behaviour.
In 2000 Sheeran and Sheina Orbell ran the same idea on cervical cancer screening with 114 women who were already equally motivated to attend. Ninety-two percent of the planners attended. Sixty-nine percent of the controls did [6]. Attendance was checked against medical records rather than asked about, which matters more than it sounds.
Then in 2002 came the study that separates the two halves cleanly. Sarah Milne, Orbell and Sheeran randomised 248 people. One intervention worked on motivation and succeeded at it, raising how threatening participants found inactivity, how capable they felt and how strongly they intended to exercise. It did not significantly raise how much they actually exercised. Adding the if-then component did [7].
You will find that study quoted everywhere with a pair of percentages attached. We are not printing them, and the reason is at the end of this article.
The Number That Went Around the World
By the mid-2000s there were enough studies to pool. In 2006 Gollwitzer and Sheeran did exactly that, gathering 94 independent tests and reporting an effect of d = 0.65 on goal attainment, which they described as medium-to-large [8].
That figure is the source of almost everything you have read about if-then plans. It became "doubles your chances", then "doubles or triples your chances", then a headline about tripling your career success. Twenty years of habit blogs and bestselling books rest on it, usually without the number appearing anywhere.
Here is the history in one view.
Notice what every popular retelling of that history leaves out. The last two lines.
Nineteen Years Later They Counted Again
In 2025 Sheeran, Olivia Listrom and Gollwitzer published a new meta-analysis. Not 94 tests. Six hundred and forty-two [9].
The raw pooled effect across all 642 was d = 0.36, with a confidence interval from 0.33 to 0.40. That is already a long way below 0.65. But the raw number is not the finding.
What they did next is the finding. Egger's regression asks whether smaller studies in a literature tend to report bigger effects, and in an honest literature there is no relationship. Here the coefficient was 1.06, which the authors note meets the published threshold for substantial publication bias. Trim-and-fill, which imputes the studies a biased literature would be missing, barely moved the estimate to 0.35. A robust Bayesian meta-analysis, which models the bias directly rather than patching around it, put the average effect at d = 0.15. The median study behind all this has about 30 people in the treatment group and 28 in the control group, ranging from 6 participants to 33995. Thirty people per arm is the typical piece of evidence behind some of the most confident advice on the internet.
Every bar in that chart is the same technique, measured by overlapping groups of researchers, using more data each time.
What a Funnel Plot Is Actually Showing You
Publication bias sounds like an accusation. It usually is not. It is the accumulated effect of thousands of ordinary decisions: a study that did not work is harder to write up, harder to get accepted, and easier to leave in a drawer.
A funnel plot is how you see it. Each dot is a study. Its horizontal position is the effect it found, and its vertical position is its standard error, which is mostly a proxy for how small it was. Biggest studies at the top, smallest at the bottom. If nothing is missing, the dots make a rough triangle, symmetrical around the true effect and spraying wider as you go down.
Here is one from a 2026 meta-analysis of if-then plans and fruit and vegetable intake.

Melum SK, Martiny-Huenger T. A systematic review and meta-analysis on the effectiveness of if-then plans - in a strict sense - to facilitate fruit and vegetable consumption in adults. Int J Behav Nutr Phys Act. 2026 Apr 15; 23:51. https://doi.org/10.1186/s12966-026-01915-y. Fig. 4. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.
Look at the bottom half. The small studies sit out to the right, at mean differences of 0.4 to 0.7. Now look at the bottom left. It is empty. No small study reports that if-then planning did nothing or made things slightly worse, even though chance alone says several should. Those studies were almost certainly run. You are looking at the shape of their absence.
That paper matters for another reason. Sanne Karlsen Melum and Torsten Martiny-Huenger accepted only interventions using a strict if-then procedure, excluding studies that used ordinary planning while calling it an implementation intention. Ten articles and twelve comparisons survived, covering 2399 people [10]. A large literature turns out to be a much smaller one when you insist that everyone actually did the thing.
Small Is Not the Same as Nothing
Stop here and you would take away the wrong conclusion.
The same 2025 analysis that found substantial bias also returned a Bayes factor of 319.3 in favour of the effect existing. That is not marginal. In Bayesian conventions it is extreme evidence. The authors are not saying if-then planning does nothing. They are saying the average is smaller than the field believed and the literature is not clean.
They also say something easy to miss, and it is the most useful sentence in the paper for anyone trying to use this. The heterogeneity statistic tau came out at 0.27, with extreme evidence that it is real.
Heterogeneity means the studies genuinely disagree with each other. Not noise. Real, systematic variation.
That changes what a small average means. If a technique produced d = 0.15 in almost everyone, it would be nearly useless. If it produces something substantial in some people and nothing in others, and you average those together, you also get 0.15. The second story fits the tau. So the honest question is not whether if-then plans work. It is what separates the cases where they do from the cases where they do not.
Here is the forest plot from the same fruit and vegetable meta-analysis, because it makes one further point that almost never gets made.

Melum SK, Martiny-Huenger T. A systematic review and meta-analysis on the effectiveness of if-then plans - in a strict sense - to facilitate fruit and vegetable consumption in adults. Int J Behav Nutr Phys Act. 2026 Apr 15; 23:51. https://doi.org/10.1186/s12966-026-01915-y. Fig. 2. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.
The pooled estimate is the black diamond: an extra 0.29 portions a day, with a confidence interval from 0.11 to 0.48. Positive, and not compatible with zero. Below it is the prediction interval, from minus 0.17 to 0.75. The confidence interval says where the average probably sits. The prediction interval says what the next study might find, and this one crosses zero comfortably.
Both are true at once. On average this works. In your case it might not.
Why the Cue Does the Work
To understand what separates the good cases from the bad ones you need the mechanism, and the mechanism is more specific than "planning helps".
An if-then plan is supposed to do two things at once. The if-clause makes the specified situation easier to notice, and the then-clause hands initiation over to that situation, so when it appears the action starts without a fresh decision.
Both halves have been tested. In 2004 Webb and Sheeran had participants form plans tied to a cue and then respond to a stream of stimuli. The planners were more accurate and faster on that cue without getting worse on the irrelevant or ambiguous ones [11]. Frank Wieber and Kai Sassenberg found in 2006 that the cue pulls attention toward it automatically, so the critical situation becomes hard to miss rather than easy to remember [12].
That distinction matters. A plan is not a note to self.
For the second half, Veronika Brandstätter, Angelika Lengfelder and Gollwitzer showed in 2001 that action initiation after an if-then plan is faster and more efficient, in the technical sense of surviving competing demands on attention [13]. In 2007 Anna-Lisa Cohen and colleagues found that if-then plans cut task-switching costs and shrank the Simon effect, both markers of executive control being used less [14].
The brain-imaging work points the same way without settling it. In 2015 Glyn Hallam and colleagues scanned 40 people regulating their emotions and found that goal intentions and implementation intentions recruited different prefrontal regions, with the plan condition leaning on areas associated with less effortful control [15].
None of that proves the action becomes automatic in the strong sense. It is consistent with the story. It is not the same as the story being true.
Where the Automatic Story Breaks Down
And there is real disagreement here, which most articles skip entirely.
Mark McDaniel and Michael Scullin published a paper in 2010 whose title is the argument: implementation intention encoding does not automatize prospective memory responding [16]. In their experiments the plan improved performance without producing the signature of a genuinely automatic process.
More recently a pair of 2023 studies pushed harder. Tim van Timmeren and colleagues used instrumental learning with outcome revaluation, the standard way of asking whether a behaviour has become habitual rather than goal-directed. If-then planning helped during acquisition and made responding more efficient early on, but it did not produce habit formation and in one design it actively impaired performance at test [17]. A companion fMRI study of the same paradigm found better early efficiency and reduced engagement of the anterior caudate, with no sign that the behaviour had gone on genuine autopilot [18].
Be careful with the older mechanism literature for a second reason. Much of it was framed around ego depletion, the idea that self-control draws on a limited resource, and a 2016 preregistered multilab replication found that effect close to zero [19]. The planning results survive. The explanation many older papers gave for them does not.
The Condition Nobody Puts in the Headline
Now the finding that reorganises everything else, and it is twenty years old.
In 2005 Sheeran, Webb and Gollwitzer looked at how implementation intentions interact with the goals underneath them. Across two studies they found a significant interaction: forming an if-then plan raised goal attainment when the person's goal intention was already strong, and did nothing when it was weak [20].
Read that again, because almost every page recommending this technique leaves it out.
The plan is not an engine. It is a transmission. It takes motivation that already exists and delivers it to a moment, and if there is nothing upstream there is nothing to deliver. Gollwitzer and Gabriele Oettingen make the same point in their 2019 overview, where the plan sits inside a wider account that begins with wanting the thing [21].
Whether you are chasing understanding or chasing a grade shapes what you do when the work gets hard, which is the subject of our piece on mastery goals versus performance goals, and whether the motivation is yours at all is the central concern of self-determination theory. An if-then plan built on a goal you do not actually hold is a sentence with nothing behind it.
I Did Not Forget. I Just Did Not Want To.
Ask people who have tried this what went wrong and the answer is consistent, and it is not the one the articles are written for. They do not say they forgot. They say the moment arrived and they did not want to.
Researchers have collected this properly. In 2022 Elina Renko, Katri Kostamo and Nelli Hankonen interviewed 19 students aged 15 to 19 at Finnish vocational schools who had all been through an intervention that taught planning, and asked why they had not planned [22]. Their answers are worth reading in their own words.
One student put the whole problem in a sentence: "I don't like to make plans myself, like if you make one and fail to follow it, it'll just make you feel bad."
Another described what making the plan felt like: "Stressful. Uncomfortable. It's like, I gotta commit."
A third: "Well, if I make it [the plan], then I get the feeling that I really need to do it, and that sometimes even brings me anxiety." The square brackets are the researchers' clarification.
That is not a failure of memory. It is a plan generating an obligation, and the obligation generating avoidance.
Older adults describe something related. In 2023 Valérie Bösch and colleagues ran think-aloud interviews with 34 people aged 65 and over while they actually formed implementation intentions [23]. Several found the exercise restrictive. "I don't plan, yeah, I just do it when I want to do it," said one participant. Others found the opposite and said so plainly: "Usually the fact that I've written something down means that I'm more likely to do it."
Both reactions come from the same study and the same task. That is the tau from the meta-analysis showing up as people talking.
So what do you do when the cue fires and you do not want to move? The honest answer is that the plan cannot fix it, because the plan was never the part supplying the wanting. If the goal is genuinely yours and the moment is just unpleasant, that is a different problem, covered in why procrastination is a mood repair strategy. If the goal is not really yours, no sentence will rescue it.
The Ways a Plan Comes Apart
Beyond motivation there are documented ways an if-then plan stops working, and knowing them beats any template. They are not five separate problems. Most are the same problem seen from different sides, which is that a cue can only carry so much before it stops being reliable.
Start with quantity. In 2012 Amy Dalton and Stephen Spiller gave participants either one everyday goal or six, and had half of each group make plans. Their result: planning helped the people carrying a single goal and reversed for the people carrying six [24]. Aukje Verhoeven and colleagues found the same shape in 2013 on one behaviour: a single implementation intention reduced unhealthy snacking and multiple plans for the same habit did not [25].
Plans compete for the same cues and the same attention, so the answer to how many you should be running is closer to one than to ten. Not because of any rule about how long a habit takes to form, which is folklore.
The same crowding turns up inside a single plan when you write it as a prohibition. Marieke Adriaanse and colleagues tested plans of the form "if I want a snack, then I will not eat crisps" against plans that named a replacement action. In that 2011 study the negation plans produced ironic effects, increasing the very behaviour they targeted [26].
If you take one practical thing from this article, take that. Never plan what you will not do. Plan what you will do instead.
Rigidity is that same reliability turning against you when the world moves. In 2012 E. J. Masicampo and Roy Baumeister gave participants a lab goal with either enough time to reach it by the planned route or not enough. With enough time the plan increased attainment. When the route was blocked it hindered people, who kept pushing at a door that was no longer open [27]. Marlone Henderson, Gollwitzer and Oettingen showed in 2007 that plans can be written for disengagement as well as for striving, which is the obvious fix and almost nobody does it [28].
This is where a real day breaks a good plan. If your cue is "when I get home at six" and you are still in the library at ten, the plan does not adapt. It just fails, and then you feel like the one who failed.
The cost of all this lands differently depending on who you are. In 2005 Theodore Powers, Richard Koestner and Raluca Topciu published a study with a title that says it: perhaps the road to hell is paved with good intentions. Among people high in socially prescribed perfectionism, the belief that others demand perfection from you, forming implementation intentions in that 2005 work was associated with worse goal progress rather than better [29]. If a plan already feels like a verdict waiting to be delivered, more plans mean more verdicts.
The machinery that makes a cue fire on time is also what makes it fire when it should not. Julie Bugg, Scullin and McDaniel found in 2013 that encoding an intention as an if-then plan doubled the rate of commission errors, which means carrying out an intention that has already been completed and should have stopped [30]. It happened in younger and older adults alike. A 2023 study of 58 students found the effect concentrated under high cognitive load [31]. Our piece on prospective memory and why you forget to do things goes further into that trade-off.
Habit strength is a moderator too. Webb, Sheeran and Aleksandra Luszczynska found in 2009 that the benefit of an if-then plan was smaller among people whose existing habit was strong, which is a polite way of saying the plans work best on behaviours that are not yet entrenched [32]. Hold that next to how habits become automatic in the first place.
What This Means If You Are Studying
Almost everything written about if-then plans uses exercise or dieting as the worked example. The studying evidence is thinner, more recent and more interesting, and it is not a set of contradictions. It is a set of conditions.
In 2021 Melinda Clark and colleagues randomised 58 Macquarie University students to mental contrasting with implementation intentions or a stress management course, and measured progress toward a self-set goal of studying more hours. The planners did better on both broad study goals at p = .038 and unit-specific ones at p = .005 [33]. The outcome was a rating of progress, not logged hours.
Now the result that complicates it, and it is the most careful study in this area. In 2025 Mirijam Schaaf, Garvin Brod and Jasmin Breitwieser ran a 42-day micro-randomised trial with 357 German fifth and sixth graders using a vocabulary app, average age 11.6. On randomly chosen evenings the children were prompted to plan the next day. Being prompted to plan did not raise the odds of studying at all, with an odds ratio of 1.07 and a p value of .39. What did matter was the quality of the plan. A child who wrote a top-tier plan was 2.41 times more likely to study than one who wrote a second-tier plan, and the bottom two tiers of plan quality were associated with worse outcomes than not being prompted at all [34].
Those participants were eleven years old, not undergraduates, and the exact figures come from the authors' preprint because the published version is paywalled. Both caveats belong with the number.
Read together with the third result, the picture holds. In 2020 Emely Hoch, Katharina Scheiter and Anne Schüler gave 119 German undergraduates implementation intentions tied to self-regulatory processes while they learned about a piano mechanism from an illustrated text. None of the plan conditions improved learning over the control group [35].
Line them up and there is no contradiction. Clark measured progress toward a behavioural goal over four weeks. Schaaf measured whether a child opened an app the next day. Hoch measured comprehension of one lesson in one sitting. Planning is a technology for getting yourself to start. Asking it to improve how well you understand a text in the next twenty minutes is asking it to do a job it was never built for.
The results that fit the technique's actual job are the strongest. Two randomised controlled trials by Elliott and colleagues tested a volitional help sheet, a menu of if-then plans matched to obstacles, against real university lecture attendance. In the 2024 trial 178 undergraduates were randomised and the planning group attended a greater proportion of lectures and kept attending for longer across an eleven-week semester [36]. The 2026 trial repeated it with 252 students on campus and asked whether rehearsing the plan improved it [37]. An earlier attendance study in 2007 found the same effect and added that conscientiousness predicted who turned up, so the plan does not work equally for everyone [38].
One more studying result is worth having. In 2010 Elizabeth Parks-Stamm, Gollwitzer and Oettingen had students take a working-memory-heavy maths exam with a television playing in the room, and tested plans aimed at shielding attention rather than at starting work [39]. Hardly anyone does this. Most people write plans for beginning a study session and none for protecting it, even though the second is where the session usually dies.
Planning also survives the meeting with distributed practice. In 2023 Breitwieser and colleagues pre-registered a study with 130 children averaging 10.8 years old in which everyone was taught why spacing study helps. Being taught it was not enough. The group that also made a plan for when and where to study was the one that kept the routine going [40]. Knowing the right schedule and running it are separate problems.
One warning from the same group. In 2024 Lea Nobbe and colleagues gave 85 students aged 10 to 12 study reminders on 16 days out of 36 and watched their vocabulary app logs. Students studied more on reminder days. The paper is titled "Smartphone-based study reminders can be a double-edged sword" [41], because an externally supplied prompt is not an internally held plan, and the first can quietly replace the second.
Writing One That Fires
Everything above narrows to a few things that decide whether your sentence does anything.
The format has to be genuinely contingent. If-then, with a real situation on the left. "I will study more this week" is a goal intention wearing a plan's clothes, and the 2025 meta-analysis found larger effects specifically when the plan used a contingent if-then structure. That is a moderator finding from the same biased literature, so read it as a direction rather than a measurement.
The cue has to be something that happens whether or not you are thinking about it. A time will do, a place is better, and an action you already perform reliably is best of all, because it brings its own reliability with it. "If I sit down at my desk after my last lecture, then I will open the problem set" has a cue you cannot miss. "If I feel motivated this evening" has no cue at all.
Cue and context are doing more work here than people expect, which is why the same plan often fires in one room and dies in another. That is the same effect described in our article on context dependent memory.
Rehearsal helps and is almost always skipped. A 2025 trial found that reinforcing an if-then plan with imagery, essentially running the moment through in your head, raised both habit strength and behaviour compared with the plan alone [42]. That is self-reported habit strength, not the laboratory sense of habit the revaluation studies tested, and the two do not have to agree. Writing the sentence once and never revisiting it produces the results people complain about.
And the plan has to be about starting. Gollwitzer and Wieber argued in 2010 that procrastination is a failure at the point of initiation rather than a deficit of desire, which is exactly what an if-then plan is shaped to fix [43]. A 2025 registered report tested daily plans against bedtime procrastination [44], and a randomised trial with 81 undergraduates tested them against academic procrastination [45]. Registered reports matter more than usual here, given what the funnel plot showed.
Habit Stacking and WOOP Are Not the Same Thing
Two related ideas get folded into this one constantly, and both deserve their own labels.
Habit stacking is the formula "after I do X, I will do Y". Every habit stack is an implementation intention with a habit as its cue. The reverse is not true, because an if-then plan can hang off a time, a place, an obstacle, a feeling or a temptation. Habit stacking is a subset, and a good one, because an existing habit is the most reliable cue you own.
Mental contrasting with implementation intentions, shortened to MCII and marketed as WOOP, is different again. It adds a step before the plan: picture the outcome you want, then deliberately picture the obstacle in the way, and only then write the if-then sentence about that obstacle. Several of the studying trials above used MCII rather than plain planning, which matters for one reason. MCII has its own effect size and it is not the if-then effect size. A 2021 meta-analysis by Guoxia Wang and colleagues pooled 24 effect sizes from 21 articles covering 15907 people and reported g = 0.336, which the authors describe as small to medium [46]. If you see that number quoted as evidence for if-then plans, someone has merged two literatures. A 2015 study found that students given MCII for time management scheduled more study hours for the week ahead than either a content control or a format control [47].
The Duckworth studies are the clearest applied cases. In 2011 Angela Duckworth and colleagues taught MCII to high school students preparing for the PSAT over a summer of self-directed study, and those students completed around 60 percent more practice questions than a placebo control [48].
The Same Technique in Eight Places
Breadth is worth having, provided every number stays attached to what it measured. Pooling these rows is the mistake this article is about.
Two of those rows carry a warning inside them. The physical activity estimate decays from g = 0.31 to 0.24 by follow-up, and the benefit sat with people whose intentions were unstable rather than those who already knew what they wanted [49]. The diet figure comes from 23 studies whose authors noticed that the better the outcome measure, the smaller the effect [50].
Both are the same warning. The better you measure and the longer you wait, the smaller the effect gets.
The children's row is the most trustworthy number in the table, because it comes from a registered report: 52 effect sizes from 42 studies covering 12957 children averaging 10.7 years old, pooling to g = 0.31 [51]. The emotion regulation row is the largest effect anywhere here at d+ = 0.91 across 21 comparisons and 1306 people, though it shrinks to medium when the comparison is a goal intention rather than no instruction [52].
The smoking and alcohol rows deserve their own sentences. In 2019 Lorna McWilliams and colleagues pooled 12 smoking cessation studies covering 15290 people and found a raw odds ratio of 1.70, then found strong funnel asymmetry, and a trim-and-fill adjustment removing three studies left OR = 0.53, which the authors describe as not significantly increasing cessation [53]. Richard Cooke, Helen McEwan and Paul Norman pooled 16 alcohol studies in 2023 and found a small reduction in weekly consumption and none at all on heavy episodic drinking [54].
Notice what smoking and drinking share. Both are behaviours people already fight hard, and both are where a plan does least.
Away from health the field-scale results are the most impressive thing here. In 2010 David Nickerson and Todd Rogers ran a voting-plan experiment across 287228 people and raised turnout by 4.1 percentage points among those contacted [55]. In 2011 Katherine Milkman and colleagues sent staff at a large firm reminders about free flu clinics, and the versions asking people to write down a date, or a date and a time, lifted vaccination above a control rate of 33.1 percent [56].
The cleanest measurement comes from medication trials, where nothing is self-reported. In 2009 Ian Brown, Sheeran and Markus Reuber randomised 81 adults with epilepsy and monitored a month of adherence with pill bottles that record every opening. The planning group took 93.4 percent of their doses against 79.1 percent, and took them on schedule 78.8 percent of the time against 55.3 percent [57].
A pattern shows through all of this. The bigger the study, the more modest the claim it supports. And every figure in that table is a raw estimate, which means most of them carry the same bias the 2025 analysis went looking for.
Not everything works. A 2009 obesity prevention trial with 709 adults found no effect on body mass index or physical activity at two weeks, three months or six months [58]. A 2014 study found people given both self-affirmation and implementation intentions were less likely to increase their exercise than people given either alone [59]. Stacking good techniques is not additive.
Those eight numbers are not comparable with one another. That is the point of the table, not a flaw in it.
Two researchers have said out loud what that spread implies. Falko Sniehotta argued in 2009 that health applications of this idea often depart substantially from the original paradigm while still claiming its evidence [60]. Martin Hagger and Luszczynska reviewed the field in 2014 and found considerable heterogeneity alongside few registered randomised trials [61].
What Survives
Strip away the marketing and a real thing is left standing.
Forming an if-then plan does something. The Bayesian evidence for a non-zero effect is extreme, and the dispute here is about size, not existence. The average is smaller than twenty years of blog posts have told you, and the researchers who supplied the original number are the ones who said so.
The average is close to meaningless on its own, because the studies genuinely disagree. Your case turns on a goal you actually hold and a cue concrete enough to fire without you watching for it. After that it is mostly hygiene. One plan on that cue rather than five. An action named, not a prohibition. And the sentence rehearsed at least once after you wrote it.
None of that is willpower advice, and that is the point. An if-then sentence helps by removing the moment where you would have had to want it enough, and replacing it with a situation you were going to meet anyway. When it fails it usually fails because one of the conditions above was missing, not because you did not try hard enough. Trying harder was never the mechanism, and what self-control actually is remains open in ways we look at in what the marshmallow test measured besides willpower.
One last thing about that 2002 exercise trial. The pair of percentages attached to it everywhere online is not in the paper. Its abstract reports no percentages at all, the full text sits behind a paywall no legitimate route reaches, and the sources quoting the control figure disagree with each other. So we left it out.
Apply that standard to us. The Christmas experiment this article opened with had 68 students in it, and the screening study had 114. Those are exactly the small studies the funnel plot tells you to be careful with. They still show most clearly what the technique does when it works, which is what small early experiments are for. They are not a measurement of how well it will work for you.
The most honest summary came from the Finnish student without meaning to. If you make one and fail to follow it, it will just make you feel bad. That is the cost of a plan built on a goal that was not yours, or hung on a cue that was never going to fire. Get those right and the sentence is cheap and worth writing. Get them wrong and it is one more thing to fail at.
Frequently Asked Questions
What is an implementation intention?
It is a plan in the form "if situation Y happens, then I will do X", which fixes in advance when, where and how you will act. Peter Gollwitzer introduced the term in 1993 and set out the evidence in a 1999 paper. The point of the format is that it links a specific cue to a specific action, so that when the cue turns up the action starts without a fresh decision.
Do if-then plans actually work?
Yes, but by less than you have been told. A 2006 meta-analysis of 94 tests reported d = 0.65, and that figure is the source of nearly every "doubles your chances" claim online. In 2025 the same research group pooled 642 tests, found a raw effect of d = 0.36, detected substantial publication bias in their own literature, and reported a bias-corrected estimate of d = 0.15. The same analysis found extreme evidence that the effect exists. So it is real, it is small on average, and it varies enormously between people and situations.
What is the difference between a goal intention and an implementation intention?
A goal intention names an outcome, like "I intend to get fit". An implementation intention names a moment, like "if it is Tuesday at seven, then I will go to the gym". The second sits on top of the first rather than replacing it, and a 2005 study found that plans raised goal attainment only when the underlying goal intention was already strong.
How do I write an if-then plan for studying?
Use a real if-then format, pick a cue that happens whether or not you are thinking about it, and name an action rather than a prohibition. An existing habit makes the best cue. Write plans for protecting a study session, not only for starting one, since that is usually where sessions die. And rehearse the sentence, because a 2025 trial found that mentally running the moment through raised both habit strength and behaviour compared with writing the plan once.
What do I do when the cue fires and I still do not want to start?
The plan cannot supply that, and no article claiming otherwise is being straight with you. Research from 2005 found that if-then plans work when the goal intention underneath is strong and not when it is weak, so a plan attached to a goal you do not really hold has nothing to deliver. If the goal is genuinely yours and the moment is simply unpleasant, that is procrastination as mood management rather than a planning failure, and it needs a different fix.
How many if-then plans should I have at once?
Fewer than you think. A 2012 study found that planning helped people pursuing a single goal and did not extend to six, and a 2013 study found that one plan reduced unhealthy snacking while multiple plans for the same habit did not. Plans compete for cues and attention. There is no evidence for any particular number of days a plan takes to settle, so ignore claims that there is.
