Introduction

Think about the last time you refreshed something. A feed, an inbox, a match screen. Most of the time there was nothing there. You did it again anyway.

You know the standard explanation. Apps borrow a trick from slot machines, unpredictable rewards are more addictive than predictable ones, and that is why your thumb keeps moving. Roughly correct. Also about one sentence deep, and the sentence is usually wrong about the machine.

Search for the phrase and you get three kinds of page. Psychology sites give you a four-cell grid: fixed or variable, ratio or interval. Therapy clinics copy the grid. Product blogs skip it and go straight to building the hook. All three assert the same fact, that a reward arriving unpredictably produces behaviour that is unusually hard to stop. Not one explains why.

There is a reason for the silence. Nobody is certain. The finding was published in 1939, called a paradox by the 1950s, and there are still at least two live theories about the cause. Experiments testing one against the other appeared in 2024 and in 2026.

This article is about that argument, and about four things the popular version gets wrong. A slot machine does not run the schedule everyone says it runs. Unpredictability does not make behaviour faster. A near-miss does not feel good. And your phone is probably not a slot machine at all.

What a Schedule Is, and Why It Was Ever Interesting

A behaviour that produces a good consequence tends to be repeated. That is the whole of the idea. What Skinner and his colleagues added is that the rule for delivering the consequence matters at least as much as the consequence itself. Feed an animal every time it presses a lever and you get one pattern. Feed it every tenth press and you get a different one. Feed it on average every tenth press, never the same number twice, and you get a third.

Those rules are the schedules, and two questions sit behind them. Does the payoff depend on how many times you act, or on how much time has passed? Is the requirement always the same, or does it move? Answer both and you have four combinations. Fixed and variable ratio count your responses. Fixed and variable interval watch the clock.

The catalogue of what each one does to behaviour was published as a book in 1957 by Charles Ferster and B. F. Skinner [1]. It runs to hundreds of pages of cumulative records, lines that climb as an animal responds. Steep line, fast responding. Flat stretch, a pause. Almost everything written since is a footnote to those shapes.

The earliest and most quoted result came nine years earlier. In 1948 Skinner reported on eight animals, all of them pigeons, fed automatically every 15 seconds no matter what they happened to be doing [2]. Six of the eight developed a repetitive ritual. One turned counter-clockwise. One pushed its head into a corner of the cage. Food had been arriving on a timer the whole time, and the birds behaved as though their own actions were producing it. Skinner called it superstition.

Nothing the birds did made any difference to when the food arrived.

That study is in every introductory textbook you will ever open. What gets left out is that it did not survive intact. In 1971 John Staddon and Virginia Simmelhag repeated the procedure while recording everything the animals did, and reported that most of the behaviour was not accidentally shaped at all [3]. It was species-typical activity landing in a predictable place between feedings. Their reading is now the standard one among researchers and almost never reaches the popular retelling.

Keep that pattern in mind as you read the rest of this. It happens more than once in this literature. A striking early result, a quieter correction, and a public memory that stops at the first one.

The Four Schedules, Side by Side

ScheduleWhat triggers the payoffWhat the behaviour looks likeWhat happens when the payoff stopsWhere you meet it
ContinuousEvery single responseSteady but unhurried. The animal knows another one is comingStops fastest of allA light switch. A vending machine that works
Fixed ratioA set number of responsesBursts of work with a long pause straight after each payoffStops quickly once the expected count passes without rewardPiecework pay. Ten coffees and the eleventh is free
Variable ratioA number of responses that changes each timeHigh and steady. Very little pausingSlow. This is the schedule that resists stoppingSales calls. Casting a fishing line
Fixed intervalThe first response after a set timeAlmost nothing early in the interval and a scallop of activity near the endStops once the expected moment passes unrewardedA bus timetable. A monthly pay cheque
Variable intervalThe first response after a changing amount of timeLow and extremely steadySlow, and steadier than variable ratioChecking whether a reply has arrived

The right-hand column is where most of the trouble starts, and we will come back to it. For now, look at the fourth column. The two variable rows are the ones that do not stop.

The Paradox Nobody Mentions

Here is the finding the whole subject rests on, and the reason it should bother you.

In 1939 Lloyd Humphreys ran an eyelid conditioning study in which one group received the air puff on every trial and another received it on a random half of the trials [4]. A light comes on, a puff of air follows, and after enough pairings the blink starts arriving at the light rather than the puff. Then he took the puff away from both groups and counted how long the blink kept coming.

The group that had been reinforced every time gave up quickly. The group that had been reinforced only half the time kept going much longer.

Read that again, because it is genuinely strange. The second group received less. Less pairing, less reinforcement, less evidence that the light meant anything. Every simple account of learning says weaker training makes weaker behaviour, and weaker behaviour should fade sooner. Instead it lasted longer.

The effect has a clumsy name. The partial reinforcement extinction effect, usually shortened to the PREE. It has been shown in rats, pigeons, mice and people, in reward learning and in fear learning. A 2024 study of conceptual fear generalisation in humans found the same asymmetry with a threat rather than a reward [5]. Another 2024 human study reported that partial reinforcement slowed acquisition and then broadened generalisation of what had been learned [6]. It is one of the most reliable findings in the whole of learning theory.

And it is a problem. Justin Harris put the difficulty plainly in 2024: the PREE defies ready explanation by associative models, because those models assume a partial schedule produces weaker conditioning that should be less resistant to extinction, not more [7]. Eighty-seven years after Humphreys, the best known phenomenon in the field is still awkward for the field's main theory.

So how do you explain it? Two ways, and they have been in tension for most of a century.

Frustration, or Memory

Abram Amsel proposed the first answer in 1958 [8]. His idea starts from something you have felt. Expect a reward, do not get it, and there is a small unpleasant jolt. Amsel called it frustrative nonreward and treated it as a real internal event with real properties.

Now run the logic. On a partial schedule the animal expects food, does not get it, feels the jolt, keeps going, and eventually does get food. The jolt has just been followed by reward, so the jolt itself becomes a cue for persistence. Frustration stops meaning stop and starts meaning keep going. When extinction arrives and the frustration is constant, the animal has already been trained to push through that feeling.

That is the frustration account. Emotion built into the habit.

Eight years later E. J. Capaldi proposed something colder [9]. Forget the feeling. What the animal remembers is the sequence. Under partial reinforcement it experiences runs of nonrewarded trials, and those runs are reliably followed by a reward. The memory of a run of nothing becomes the signal that something is due. Extinction is one long run of nothing, which is precisely the cue the animal was trained on.

That is colder than Amsel's account and a great deal easier to test.

Sequential theory is the more popular account today, and it makes a sharp prediction: the length of the nonrewarded runs during training should determine how long the animal persists during extinction. In 2024 Harris tested exactly that and found the sequencing of trials during partial reinforcement did shape later extinction [7].

Then it got complicated.

The Result That Made It Complicated

In 2026 Mirari Wilcher and Justin Harris trained four groups of rats with both a consistently reinforced stimulus and a partially reinforced one, manipulating the timing and the sequencing of reward separately [10]. The design separated two things that normally travel together: whether the animal tracks a run of trials, or simply tracks elapsed time.

The awkward part had already appeared in earlier work from the same group. Rats that learned to anticipate reward after a run of nonreinforced trials showed a smaller partial reinforcement extinction effect, not a larger one. If the memory of the run is what carries persistence, making that memory more useful should not weaken the effect.

A 2022 result points the same way. Shanae Norton and Harris found that waiting before you start withholding reward shrinks the persistence advantage rather than preserving it [11], which is hard to square with a pure sequence-memory account and easier to square with something that fades.

Two further explanations are on the table. A 2017 neurocomputational model builds the effect out of two interacting processes, one associative and one affective, which is Amsel with the arithmetic filled in [12]. Patrick Anselme argued in 2021 that these paradoxes dissolve if effort is treated as motivated by uncertainty rather than merely spent on it [13].

So you are left with a real, unresolved question, which is a more honest place to stand than the one the rest of the internet offers. Something about inconsistent reward makes behaviour durable. Whether that something is a repurposed emotion, a memory of runs, or a timing signal is not decided.

1939
Humphreys shows partially reinforced blinks outlast consistently reinforced ones
1948
Skinner reports superstitious rituals in eight pigeons fed on a timer
1957
Ferster and Skinner publish the catalogue of schedule performances
1958
Amsel explains persistence through frustration that becomes a cue
1966
Capaldi explains it instead through memory for runs of nonreward
1971
Staddon and Simmelhag reinterpret the superstition result
1977
Matthews and colleagues show uninstructed humans behave like animals
1980
Hurlburt and Knapp find no behavioural gap between variable and random ratio
1983
Mazur locates the schedule effect in pausing rather than in speed
1990
Rachlin reframes persistent gambling as a pattern of choice
1997
Schultz Dayan and Montague identify the reward prediction error signal
2003
Fiorillo Tobler and Schultz find a second signal that tracks uncertainty
2009
Clark shows near-misses are unpleasant and still increase the urge to play
2010
Dixon documents losses disguised as wins on multi-line machines
2024
Palmer and colleagues replicate the near-miss effect under preregistration
2026
Wilcher and Harris report evidence that troubles sequential theory

Notice how much of that timeline is correction rather than discovery. This field keeps revisiting its own results.

What Actually Changes Is the Pause

The mechanism is smaller and more concrete than the story. That story says unpredictable rewards send the brain into a frenzy of seeking, and if it were true you would expect the behaviour to speed up. In 1983 James Mazur checked, running three groups of rats on ratio schedules matched for average effort [14]. The animals pressed a lever for milk. One group was compared on fixed against mixed ratios, one on fixed against random ratios, and a third worked on variable-interval schedules for contrast.

The result is the quiet centre of this whole subject. On every ratio schedule, fixed or random, the gaps between presses were distributed the same way, peaking at roughly 0.2 seconds [14]. The rats were not pressing faster under uncertainty. They were pressing at the speed rats always press.

What differed was pausing. After a payoff on a fixed ratio, a rat stops for a while before starting the next run of work. Mazur found those pauses were reliably longer on fixed-ratio schedules than on mixed-ratio or random-ratio ones. Uncertainty does not add energy. It removes the rest.

That changes what the effect feels like from the inside. On a predictable schedule there is a moment after each reward when you know the next one is far away. That moment is where you stop. It is where you put the phone down, cash out, close the tab. Make the schedule unpredictable and the moment never arrives, because the next payoff could be the next attempt.

You do not keep going because unpredictability is thrilling. You keep going because it never gives you a good reason to stop.

Fixed ratio

Random ratio

Payoff arrives

Is the next one predictable?

Long pause

No pause

A natural stopping point

Straight back to responding

There is a related idea worth naming here. John Nevin and Randolph Grace argued in 2000 that resistance to change is a separate property of behaviour from how fast it happens, with its own determinants [15]. They called it behavioural momentum. How hard a behaviour is to disturb is not the same question as how vigorous it looks, and schedules pull the two apart.

The Machine Is Never Due

This is where every page about this subject makes the same mistake.

Open any explainer and you will read that a slot machine runs on a variable ratio schedule. It is the flagship example, on the psychology sites, the therapy blogs, the encyclopedias and the growth-marketing posts.

It is the wrong schedule.

A variable ratio schedule, in the technical sense, has a predetermined series of response requirements averaging out to some ratio. Ten presses, then three, then twenty-two, then five, worked out in advance. A gaming machine does not do that. It generates each spin afresh from a random number generator, with no memory of what came before. That is a random ratio schedule, and gambling researchers have distinguished the two for decades. In 2023 Paul Delfabbro, Daniel King and Jonathan Parke stated it directly in a review: most forms of gambling, and slot machine play in particular, follow a random ratio schedule [16].

Why does the distinction matter. Because of what a player can infer.

Under a genuine variable ratio there is a fixed pool of outcomes being worked through, so a long run of losses really does mean the remaining payoffs are getting closer. Under a random ratio there is no pool. A machine that has not paid in four hours is in the same state as one that paid thirty seconds ago. The feeling that it is due is not an exaggeration of something true. It describes a mechanism that is not there.

John Haw made a related point in 2008 [17]. Comparisons between gambling schedules usually rest on the average frequency of wins, and Haw argued from machine data that this is the wrong statistic. What predicts persistence is where the wins fall: how many arrive early, and how long the runs of nothing get. Two machines with the same advertised payout can feel entirely different.

Now the honest part. If the two schedules really are psychologically different, somebody should be able to measure a difference in behaviour. In 1980 Russell Hurlburt and Terry Knapp did the obvious experiment, putting 20 subjects on a simulated slot machine task with variable ratio and random ratio schedules available concurrently [18]. The subjects did not behave differently. Not in which game they chose, and not in the strategy they used.

Twenty people in 1980 does not close the question. It does set the limit of what this article can claim. The distinction corrects how the schedule is described and what a gambler can conclude from a losing run. It is not a measured gap in behaviour, and anyone who says it is what makes slot machines dangerous has gone past the evidence.

What is not in doubt is that the machines are engineered. In 2009 Kevin Harrigan and Mike Dixon obtained the design documents for slot machines used in Ontario through freedom-of-information law [19]. Those PAR sheets specify speed of play, stop buttons, bonus modes, nudges, near misses, and the fact that two machines that look identical on the floor can pay back very differently. Harrigan had already shown in 2008 that near misses are manufactured, by loading high-paying symbols onto the virtual reels so they land next to the pay line far more often than a fair reel allows [20].

Worth pausing on that. Most consumer products do not have public design documents.

Mark Griffiths made the case in 1993 for studying these structural features rather than the gambler [21]. Three decades later that is the mainstream position.

One more framing is worth keeping. Howard Rachlin argued in 1990 that persistent gambling despite heavy losses is not necessarily irrational at the level of the single bet, because people do not choose bets one at a time [22]. They choose patterns that extend over time, and inside a pattern that ends in a win the individual losses look different. Many disagree. It is still a useful antidote to the assumption that a persistent gambler is failing at arithmetic.

The Near-Miss Feels Bad. You Play On Anyway.

This is the part that changed while this article was being researched.

The going-in assumption was the one you have probably absorbed. A near-miss feels almost as good as a win, so the brain treats it as a partial reward and you keep going. Tidy story. Not what the data show.

In 2009 Luke Clark and colleagues ran a simplified slot machine task with 40 healthy volunteers rating each outcome as it happened, and repeated it with a separate 15 volunteers in a scanner [23]. Near-misses were rated as less pleasant than full misses. Less pleasant. And the same outcomes increased the rated desire to keep playing.

Both at once. The outcome that feels worst is the one that makes you continue.

A second condition mattered as much. The effect appeared only on trials where the participant had chosen their own gamble rather than having it assigned. Personal control, over an outcome control cannot possibly influence, was what switched it on. In the scanner, near-misses recruited striatal and insular regions that also responded to real wins.

The result has held up in its general shape and has been sharpened in its details. In 2010 Henry Chase and Clark tested 20 adults who gambled anywhere from recreationally to severely, and found that midbrain response to near-misses scaled with gambling severity [24]. In 2016 Guillaume Sescousse and colleagues scanned 22 patients with gambling disorder alongside healthy controls and reported amplified striatal responses to the same outcomes [25].

Then came the result that turns a correlation into a mechanism. In 2014 Clark and colleagues took the question to patients with focal brain lesions and found that damage to the insula abolished the gambling distortions that near-misses normally produce [26]. Not reduced. Abolished. Take out one region and the effect is gone.

That is a strong result built on a small number of very rare patients.

Then, in 2024, the correction that every popular account skips. Lucas Palmer, Mario Ferrari and Luke Clark ran preregistered conceptual replications with 169 and 148 participants rating outcomes and a further 170 participants measured on behaviour, using an online slot simulator that made about a third of spins near-misses [27]. Those samples are several times the size of the 2009 original, and the picture that came back was more limited. Clark is a co-author of both, which is what a healthy literature looks like.

So the accurate summary is this. Near-misses are unpleasant, they raise the reported urge to continue, they depend on the illusion that you are steering, and their effect on behaviour is smaller and more conditional than the first wave suggested.

The effect is also not confined to gambling. In 2017 Chanel Larche and colleagues recruited 60 participants who played Candy Crush regularly and recorded heart rate, skin conductance, arousal, frustration and urge to play across 30 minutes of real play [28], sorting every outcome into levelling up, failing clearly, or just missing the level. No money was involved anywhere in the design.

Two of that paper's figures tell the story better than a paragraph can.

Bar chart of mean subjective frustration ratings for wins losses and near-misses

Larche CJ, Musielak N, Dixon MJ. The Candy Crush Sweet Tooth: How 'Near-misses' in Candy Crush Increase Frustration, and the Urge to Continue Gameplay. J Gambl Stud. 2017 Jul 20; 33(2):599-615. https://doi.org/10.1007/s10899-016-9633-7. Fig. 7. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.

Frustration was rated after each outcome on a scale running from 1 to 7, and near-misses came out highest of the three outcome types, significantly above plain losses [28]. Wins sat near the bottom of that scale and both kinds of failure well above the middle. Just missing was worse than missing.

Bar chart of mean urge-to-play ratings for wins losses and near-misses

Larche CJ, Musielak N, Dixon MJ. The Candy Crush Sweet Tooth: How 'Near-misses' in Candy Crush Increase Frustration, and the Urge to Continue Gameplay. J Gambl Stud. 2017 Jul 20; 33(2):599-615. https://doi.org/10.1007/s10899-016-9633-7. Fig. 8. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.

The urge chart tells the other half. Near-misses produced significantly more urge to keep playing than plain losses did, while wins and near-misses did not differ significantly [28]. Read that second figure carefully. Its axis starts at seven rather than at the bottom of the scale, and urge was scored by summing two items running from 1 to 7, so the real range is two to fourteen and every bar sits near the middle of it. The frustration gap is the large one. The urge gap is real and visually exaggerated by where the axis begins.

The same paradox, in a puzzle game with no stakes. The outcome people liked least kept them playing.

A Win That Is Really a Loss

Modern machines have a second trick, arguably more effective than the near-miss because most players never notice it.

On a multi-line slot you bet on many lines at once. Land a winning combination on one and the machine lights up, plays the win sound, and credits you less than you staked on that spin. You have lost money. Everything about the presentation says you won. Researchers call these losses disguised as wins.

In 2010 Mike Dixon and colleagues measured skin conductance and heart rate in 40 people new to slot machines as they played a multi-line game [29]. The physiological response to a loss disguised as a win matched the response to a real win, and both were significantly larger than the response to a plain loss. The body was not doing the arithmetic.

The sounds are doing the work. A follow-up study put 157 participants through 300 spins each with the usual celebratory audio, with silence, or with a negative sound on those outcomes, and used the audio to expose them as the losses they are [30]. In 2019 Candice Graydon and colleagues randomised gamblers to games with few, moderate or many of these outcomes, then let them keep playing for as long as they liked through a losing streak they did not know had been arranged [31]. Persistence differed by condition and interacted with problem-gambling symptoms. And a 2023 study found that these outcomes contribute to players overestimating how much they had won [32].

Put the two together and a modern machine is not really running a schedule of wins at all. It is running a schedule of win-shaped events, only some of which are wins. The reinforcement is the sight and the sound, not the money.

The Dopamine Story, Corrected

You have been told that variable rewards flood the brain with dopamine. The real finding is stranger and better.

Start with what is settled. In 1997 Wolfram Schultz, Peter Dayan and P. Read Montague described the signal that midbrain dopamine neurons actually carry [33]. It is not pleasure and it is not reward. It is a prediction error. Better than expected produces a burst. Exactly as expected produces nothing at all. Worse than expected produces a dip below baseline. A reward you fully saw coming barely moves the signal.

That reframes unpredictable rewards. A predictable payoff, once learned, stops generating a signal. An unpredictable one never stops, because it can never be fully predicted.

Then, in 2003, Christopher Fiorillo, Philippe Tobler and Schultz found something they had not been looking for [34]. Recording from dopamine neurons in monkeys while cues signalled different probabilities of reward, they saw the expected prediction-error burst varying with probability. They also saw a second, previously unobserved response: a gradual rise in activity that built through the delay and peaked at the moment reward was possible. That ramp did not track how much reward was coming. It tracked uncertainty. Uncertainty is greatest when the odds are even, so the ramp was largest at p = 0.5 [34].

Fifty-fifty is the worst case for prediction and the biggest case for that signal.

The interpretation was contested and Fiorillo defended it, reporting in 2011 that midbrain dopamine neurons are transiently activated by reward risk itself rather than by the reward it might deliver [35]. By 2008 Schultz and colleagues had found explicit uncertainty signals in people, so this is not a monkey-only result [36].

The cells that do this are not only in the midbrain. In 2013 Ilya Monosov and Okihide Hikosaka found neurons in the monkey anterodorsal striatum that code uncertainty in a graded way rather than all at once [37].

There is a further wrinkle that speaks directly to refreshing a feed. In 2021 Ahmad Jezzini and colleagues found a prefrontal network that handles the preference for advance information about an uncertain outcome [38]. Wanting to know is its own motivation.

The last correction is the one that matters most for how you read all of this. Dopamine is not the pleasure chemical. Terry Robinson and Kent Berridge have spent three decades arguing that mesolimbic dopamine carries wanting rather than liking, and their 2025 review sets out where the theory stands after thirty years [39]. Wanting can be amplified without any increase in enjoyment. That is exactly the shape of the near-miss result, and it is exactly the shape of the feeling you have when you check something for the ninth time without expecting anything.

Uncertainty appears to do this on its own, without any drug involved. In 2015 Mike Robinson and colleagues found that exposing rats to uncertain reward raised incentive salience about as much as sensitising them with amphetamine did [40]. In 2017 Fiona Zeeb and colleagues reported that the same exposure left rats behaviourally sensitised and choosing the risky option more often afterwards [41].

Patrick Anselme and Mike Robinson had proposed the reading in 2013: dopamine's part in gambling makes more sense as motivation for an uncertain reward than as pleasure at getting one [42]. Martin Zack and colleagues built a full account of gambling addiction on that signal in 2020 [43].

The word to retire is flood. What uncertainty does is keep a signal alive that a predictable reward would have switched off. Our companion piece on dopamine and learning goes further into what that molecule is actually for, and the article on predictive processing covers the wider machinery that prediction error belongs to.

Ratio Is Not Interval, and Your Phone Knows the Difference

Back to the four-cell grid, because there is one distinction inside it that does more work than everything else on this page.

A ratio schedule pays off according to how many times you act, so acting twice as often pays you roughly twice as often. An interval schedule pays off according to how much time has passed, on the first response after the clock runs out. Act twice as often there and you get almost nothing extra, because the reward was waiting for the clock and not for you.

That difference produces a reliable difference in behaviour, and it holds in humans. In 1977 Byron Matthews, Eliot Shimoff, A. Charles Catania and Terje Sagvolden worked with two students at a time, in separate rooms, in what is still the cleanest demonstration of it [44]. Each student pressed a telegraph key. Presses were occasionally reinforced with a light, during which button presses earned points exchangeable for money. The instructions described the button and said nothing about the telegraph key, so key pressing had to be shaped, exactly as it would be in an animal.

One of the pair was on a variable ratio schedule and the other on a variable interval schedule yoked to it, so both received reinforcement at the same moments and only the rule producing it differed.

The variable ratio students pressed faster. With some pairs the effect appeared so quickly that reversing the schedules reversed the rates within a single session. Uninstructed humans behave like pigeons.

Now the part that matters more. With other pairs the sensitivity was not reliable, and the difference lay in how the behaviour had started. When an experimenter demonstrated the key press instead of letting it be shaped, the schedule stopped controlling the rate cleanly. Being shown what to do inserts a rule between the person and the contingency, and the rule wins.

This is why you should be careful with any confident claim mapping an animal schedule onto a human product. Delfabbro, King and Parke set out the caveats in 2023 [16]. People bring verbal rules. They form appraisals about what the contingency is. They modify the schedule themselves, choosing the stake, the speed, the number of lines. None of that exists for a rat, and all of it changes the outcome.

With that caution in place, here is the reframing.

SituationClosest scheduleDoes acting more often get you more?
A slot machine spinRandom ratioYes. Every spin is an independent chance
Cold calling for salesVariable ratioYes. More calls means more sales on average
Refreshing a social feedVariable intervalBarely. New posts arrive on other people's time
Checking your inboxVariable intervalNo. The email arrives when it arrives
Waiting for a reply to a messageVariable intervalNo. Checking cannot make anyone answer

Look at the right-hand column. Three of those five buy nothing extra, and they are the three most people do most often.

That is not a small correction. If your feed were a slot machine, refreshing more would produce more reward and the behaviour would at least be rational. It is not. Something new appears because somebody else posted it, a function of elapsed time, and your checking has almost no influence on the supply. What it guarantees is that whenever something arrives, you are first to know.

Variable interval schedules produce exactly what you would expect from that. Low, steady, persistent responding. Not frantic. Just constant, and very hard to stop.

The empirical work on phones fits. In 2012 Antti Oulasvirta and colleagues showed that smartphone use is built out of brief repeated checks lasting seconds rather than long sessions, and that those checking habits are what make the device pervasive rather than any single application [45]. Christian Montag and colleagues argued in 2019 that the addictive potential sits in specific design features of social platforms and freemium games rather than in the medium as a whole [46].

Interventions that target the alert rather than the schedule have a mixed record, which is worth knowing before you install anything. A 2020 study found that pop-up notifications telling people about their own screen time had limited effect on how often they actually checked [47], and a 2025 experiment that disabled notifications outright reported mixed effects rather than a clean reduction in use [48]. Turning off the alerts changes the cue. It does not change the schedule underneath, which is that things keep arriving whether you are told about them or not.

A schedule-driven behaviour is not the same as one you decided on, and our article on how habits become automatic covers that transition. Mark Bouton showed in 2021 that the switch between the two is governed by context far more than by intention, which is why the same person behaves differently in two rooms [49].

Why Just Stopping Does Not Work

If you have ever tried to quit a checking habit by deciding to, you know what happens next. It gets worse before it gets anything.

The extinction burst is the temporary surge in responding that follows the removal of reinforcement. Press the lever, nothing happens, press harder and faster. It feels like proof that the habit is winning.

Here is the finding that should change how you think about it. In 2025 Timothy Shahan and Matias Avellaneda trained rats to press a lever for a single food pellet on a variable-interval schedule averaging 1.5 seconds, then switched them to extinction with no alternative reward available, or with one pellet available, or with six [50]. A large extinction burst appeared when there was nothing else available. Both how often it happened and how big it got fell when there was.

The burst is not a law of nature. It is what happens when you take a source of reward away and put nothing in its place. Work published in 2026 shows it can be shaped, made more likely or prevented outright, by how the transition is arranged [51].

Eliminated behaviour also comes back. Resurgence is the return of an extinguished response when the alternative that replaced it weakens, and in 2021 Shahan and Brian Greer reanalysed two clinical studies and found the size of the return tracked the size of the cut in alternative reinforcement [52].

Neither finding is a motivational slogan. What predicts whether a schedule-maintained behaviour stays gone is whether something else is reliably paying. Not willpower. Not insight. Availability of an alternative.

That is roughly what people doing it describe. In 2024 Emily Nolan and colleagues interviewed thirty participants in Quebec who gambled anywhere from recreationally to problematically, and asked how they actually try to hold themselves in check [53]. Very little of what came back was resolve. Most of it was arrangement. One participant put it this way: “I’ll put a lot of obstacles in front of me for my money.” Another had stopped carrying a debit card and brought a fixed amount of cash instead. The paper reports that participants scoring higher on problem-gambling severity found these strategies hardest to keep up. That is the same finding from the other side: what fails is the arrangement, not the intention.

The intervention evidence says something similar. A 2021 systematic review and meta-analysis pooled the trials of responsible-gambling pop-up messages and found effects on both behaviour and cognition [54]. Messages are not nothing. They are also not a schedule change.

And it cuts both ways, which is the fairest thing this article can say about the mechanism. In 1980 Jack Nation and Donald Woods argued that partial reinforcement belongs in psychotherapy on purpose, because it builds behaviour that survives the stretches when nothing is going well [55]. If a study habit only ever pays off in an obvious way, it will stop the first week it does not. If the payoff has always been irregular, an unrewarding week is just a week. Our pieces on delayed gratification and on gamification reach the same design question from different directions.

There is a warning attached. Adding rewards to something a person already wants to do can reduce the wanting, which is the central finding of self-determination theory, and it means the schedule is a tool for behaviour you are trying to sustain against friction, not for behaviour that was fine on its own.

What This Article Is Not Claiming

A few limits, stated plainly, because the rest of the internet skips them.

Most of the schedule literature is animals. Mazur's pauses are rats. Skinner's superstition is pigeons. Shahan's extinction burst is rats. Where a finding has been shown in humans this article said so, and where it has not it did not pretend otherwise.

The near-miss literature is mostly small and mostly laboratory. Forty people, fifteen, twenty, twenty-two. The 2024 replication series with samples of 148 to 170 is the largest test here, and it reported a more limited picture.

The random ratio point corrects a description, not a measured behavioural difference. Hurlburt and Knapp's twenty subjects found no gap. Run that comparison again with modern methods and this section may need rewriting.

And nothing here is about you specifically. Gambling disorder is a clinical literature about a diagnosed condition, and a 2023 systematic review of craving in it covers a population this article is not describing [56]. Checking a phone is not a disorder. If you want the boundary between a habit and something that has stopped being one, our article on addiction and the learning system is where that argument lives.

What To Take From This

The mechanism is simpler than the mythology and more interesting than the summary.

A reward you cannot predict never lets you conclude that the next one is far away. That removes the pause, and the pause is where stopping happens. Rats on a random ratio do not press faster than rats on a fixed ratio. They just never rest.

Why that makes behaviour so durable is still argued over. Frustration turned into a cue, memory for runs of nothing, a timing signal, or some mix of them. Amsel and Capaldi set out the positions in 1958 and 1966 and experiments in 2024 and 2026 are still testing them.

The examples everyone uses are half wrong. The slot machine is a random ratio, so it is never due. The near-miss is unpleasant and motivating at once. The feed in your pocket is closer to a variable interval, so the extra checks are not buying the extra chances you think they are.

And the way out is not a decision. It is a different schedule, with something else reliably paying.

Frequently Asked Questions

What is a variable reward schedule?

A variable reward schedule is a rule for delivering reinforcement in which the requirement changes from one payoff to the next, so you cannot predict when the next reward will arrive. There are two kinds and the difference matters more than most explanations admit. On a variable ratio schedule the payoff depends on an unpredictable number of responses, as in cold calling. On a variable interval schedule it depends on an unpredictable amount of elapsed time, as in checking for a reply. On the interval version, responding more often does not earn you more. The catalogue of what each schedule does to behaviour was published by Charles Ferster and B. F. Skinner in 1957.

Why are unpredictable rewards harder to quit than predictable ones?

The behavioural fact has been established since 1939, when Lloyd Humphreys showed that human eyelid responses reinforced on only half of trials kept going far longer once reinforcement stopped. That is the partial reinforcement extinction effect. The mechanism is still argued over. Abram Amsel proposed in 1958 that the frustration of an expected reward failing to appear becomes a cue for persistence, because on a partial schedule that frustration is repeatedly followed by reward. E. J. Capaldi proposed in 1966 that animals instead remember runs of nonrewarded trials and treat the memory of a run as a signal that reward is due. Experiments published in 2024 and 2026 are still testing the two accounts, and some recent rat results sit awkwardly with the sequential explanation.

Are slot machines really a variable ratio schedule?

No, and almost every explanation says otherwise. A variable ratio schedule works through a predetermined series of response requirements that average out to a ratio. A modern gaming machine generates each outcome afresh from a random number generator with no memory of previous spins, which is a random ratio schedule. The practical consequence is that a machine is never due, because there is no pool of outcomes being worked through. One caution belongs with this. When Russell Hurlburt and Terry Knapp compared the two schedules directly in a simulated slot machine task with 20 subjects in 1980, they found no behavioural difference between them.

Do near-misses really make people keep playing?

They increase the reported urge to continue, and they do it while feeling worse rather than better. In 2009 Luke Clark and colleagues found that 40 healthy volunteers rated near-misses as less pleasant than plain losses while reporting a greater desire to keep playing, but only on trials where they had chosen their own gamble. The effect scales with gambling severity and is amplified in patients with gambling disorder, and damage to the insula abolishes it. A preregistered replication series in 2024 with samples of 148 to 170 participants reported a more limited picture, so the honest summary is that the effect is real, conditional and smaller than first advertised. It is not confined to gambling either. Among 60 Candy Crush players with no money at stake, near-misses were the most frustrating outcome and produced significantly more urge to continue than plain losses.

How do you break a habit that runs on unpredictable rewards?

The research points at the schedule rather than at resolve. In 2025 Timothy Shahan and Matias Avellaneda found that the extinction burst, the surge of responding that follows the removal of reward, appeared clearly in rats only when no alternative reward was available, and shrank when one was. In 2021 Shahan and Brian Greer reanalysed two clinical studies and found that eliminated behaviour returned in proportion to how much the replacement reinforcement had been cut. Both point at the same conclusion, which is that whether the behaviour stays gone depends on whether something else is reliably paying. Most of this evidence is animal work or clinical work with children, so treat it as a direction rather than a protocol.