Introduction
Imagine you have been sitting at a table for an hour.
In front of you is a wooden tray and twelve spools. Your job is to put the spools on the tray, one at a time, using one hand. When the tray is full you take them all off. Then you put them back on. You do this for thirty minutes.
Then the tray is taken away and replaced with a board holding forty-eight square pegs. Your job now is to turn each peg a quarter turn clockwise. Then turn it another quarter turn. Then another. You work along the board, peg by peg, with one hand, for another thirty minutes.
Nobody tells you why. There is no puzzle to solve, no score, no end point you are working toward. There is just the next spool and the next quarter turn, and a man with a stopwatch who does not explain himself.
That hour is the setup for one of the most quoted experiments in the history of psychology, and almost nobody who quotes it knows what happened in it.
Here is what happened. When the hour ended, some of the men were asked to do one more thing. Go into the waiting room, they were told, and tell the next participant that the task was interesting and enjoyable. Some were offered one dollar to say it. Others were offered twenty.
Afterwards, everyone was asked how much they had actually enjoyed the task.
The men who had been paid one dollar said they had enjoyed it. The men who had been paid twenty dollars said they had not. Paying somebody twenty times more to say a thing left them less likely to mean it, not more.
That result, published by Leon Festinger and James Carlsmith in 1959, is the load-bearing evidence for cognitive dissonance [1]. It is on the syllabus of essentially every introductory psychology course in the English-speaking world. It has been retold so many times that the retellings now cite each other rather than the paper. That is what happens to famous experiments, and the same drift is documented in the popular account of what the Bobo doll studies really showed.
And in 2024, thirty-nine laboratories in nineteen countries ran four thousand eight hundred and ninety-eight people through the same family of experiment, and the effect that is supposed to be doing the work did not appear [2].
This article is about what really happened in that hour in 1959, what the numbers actually were, and what sixty-seven years of argument have done to them. It is not a story about a theory being proved. It is not a story about a theory being demolished either, and if you have seen that headline somewhere, it was wrong. It is a story about an idea whose evidence turned out not to be where almost everybody thinks it is, and about what is still standing once you go and look.

The Sentence That Started It
Two years before the pegs, Festinger published a book called A Theory of Cognitive Dissonance (Stanford University Press, 1957). The argument in it is short enough to fit in a sentence.
People do not tolerate holding two thoughts that contradict each other. The contradiction is not merely noticed. It is uncomfortable, in something close to the way hunger is uncomfortable, and that discomfort pushes the person to get rid of it.
That is the whole theory. Everything else is detail about what counts as a contradiction and which way people jump to resolve it.
What made it interesting in 1957 was the second half. If the discomfort is a drive, then it does not much care how you switch it off. A person who has done something that contradicts what they believe has two obvious routes available. They can undo the behaviour, which is often impossible because it has already happened. Or they can move the belief.
Moving the belief is cheap. And crucially, the person doing it does not experience it as moving a belief. They experience it as noticing something true.
Festinger had already spent time watching this happen outside a laboratory. In 1956 he and two colleagues published When Prophecy Fails (University of Minnesota Press), an account of a small group who expected the world to end on a specific date and had made costly, public, irreversible commitments on that basis. When the date passed and nothing happened, the members did not conclude they had been wrong. Several of them became more committed, not less, and started recruiting.
That book is often described as the origin of the experiment, which overstates the connection. The theory was published separately, and the 1959 study is a laboratory test built from scratch rather than a follow-up. But the book explains the shape of the intuition Festinger was working from. He had watched people who could not take an action back, and he had watched what happened to their beliefs instead.
The problem with a field observation is that you cannot run it twice. So the question became how to make a person do something that contradicted what they thought, on purpose, in a room, with a control group.
The answer was to make them lie about spools.
What the Sixty Men Actually Did
The details matter here, because almost every popular retelling strips them out, and some of what gets stripped out is the most interesting part.
Seventy-one male undergraduates from an introductory psychology course at Stanford were recruited [1]. They came in believing they were taking part in a study of performance measures.
They did the hour. Spools for thirty minutes, pegs for thirty minutes, one hand, no explanation.
Then the experimenter told a third of them that the study was over and walked them to an interview. That was the control group. They had done the boring hour and nothing else.
The other two thirds got a different story. The experimenter explained, apologetically, that the study compared people who came in cold against people who had been told in advance that the task was enjoyable. The student who normally did the telling had not shown up. Could they do it, just this once. There would be a payment.
For one group the payment was one dollar. For the other it was twenty.
Almost everyone agreed. They went to the waiting room, met a young woman they believed was the next participant, and told her the hour ahead of her would be interesting and fun. She said something noncommittal about a friend who had said it was boring. They reassured her. Then they were taken to a separate interview, run by someone they were told had nothing to do with the experiment, and asked what they had thought of it.
Laid out as a sequence, the procedure is a small piece of theatre with four parts.
The money sits early in that sequence. It arrives before the lie and is never mentioned again. By the time the last arrow is drawn, the payment is a fact about the past, and the only thing left in the room is a man and an opinion he now has to give.
Eleven of the seventy-one never made it into the results. Five worked out that something was off about the payment. Two told the woman in the waiting room it was a set-up. Three refused to take the money at all. One would not proceed until he had her phone number.
That leaves sixty men. Twenty in each condition.
Twenty. Hold on to that number, because it is going to matter later, and because you will not find it in most of the pages that describe this experiment as a landmark.
The Four Questions, and the Three That Did Nothing
Here is the part that gets lost.
The interview did not ask one question. It asked four. How enjoyable was the task. How much did you learn from it. How scientifically important was it. Would you take part in a similar experiment again.
Only one of them separated the groups.
Look at the enjoyment row first, because that is the famous one. The control group, who did the hour and told nobody anything, rated it at minus 0.45. Mildly unpleasant, which is a fair description of an hour of spools. The men paid twenty dollars rated it at minus 0.05. Statistically that is sitting next to the control group. Being paid well to say it was fun did essentially nothing to what they thought of it.
The men paid one dollar rated it at plus 1.35.
Against the control group that difference gives a t of 2.48, with a p of .02. Against the twenty dollar group, a t of 2.22 and a p of .03 [1].
Now look at the other three rows, where almost nothing happens. Asked how much they had learned, the three groups came in at 5.60, 6.45 and 6.03, which is noise wearing a decimal point. On scientific importance the spread is barely wider. Only the last question, whether they would come back and do the whole thing again, moves in the same pattern as enjoyment, and it moves less.
Three of the four measures did close to nothing. The entire famous result is one row of that table, a gap of about 1.4 points on an eleven-point scale, between two groups of twenty men.
That is not a criticism. It is a real, statistically detectable effect and it went in the direction nobody expected. But it is a very different object from the one that gets described in most retellings, where a bold prediction meets an overwhelming confirmation. What actually happened is a modest effect on one dependent measure that pointed the opposite way to the prevailing theory of the time, which is precisely why it mattered so much.
The people who most need to know this are the ones most confident about the study. That gap between confidence and detail is a pattern in itself, and it is the subject of a separate piece on the illusion of knowing.
What Twenty Dollars Was
There is a second detail that almost every retelling flattens, and it changes how the whole thing feels.
Read as modern money, one dollar and twenty dollars are both small. One is nothing. Twenty is a sandwich. If you read the experiment that way, both groups look like they were being asked to lie for pocket change, and the difference between them looks trivial.
In 1959 it was not trivial. Using the long-run United States consumer price index, prices have risen by roughly a factor of ten or eleven since then, so twenty dollars in 1959 buys something in the region of two hundred dollars today, and one dollar buys around ten. Treat those as approximations. The ratio you get depends on which index and which end date you use, and no single figure is defensible enough to put in print.
Approximate is enough. Twenty dollars was a serious sum to hand a student for ten minutes of talking. One dollar was almost insulting.
That is the engine of the whole result. The man paid two hundred dollars in today's money has a complete and satisfying explanation for what he just did. He lied because he was paid well. There is no contradiction to resolve, because the money resolves it.
The man paid ten dollars in today's money has no such story. He lied to a stranger, looked her in the eye, and told her that an hour of turning pegs was interesting. And the payment is not big enough to explain it. So he is left holding two thoughts that do not sit together: I am a reasonable and honest person, and I just told a stranger something false for almost nothing.
One of those has to move. The behaviour cannot. So the belief does.
By the time he sits down for the interview, the hour has quietly become more interesting than it was.
The Prediction That Ran Backwards
To see why this landed the way it did, you have to know what it was arguing against.
The dominant view of attitude change in the 1950s came out of reinforcement thinking. Reward a behaviour and you strengthen whatever goes with it. On that account, paying somebody more to say something should make them more likely to believe it, not less. More reward, more learning, more change.
Festinger and Carlsmith predicted the opposite, and stated it plainly. Their paper sets out two derivations. First, that a person induced to say something contrary to a private opinion will tend to shift the opinion toward what was said. Second, and this is the sharp one, that the larger the pressure used to produce the statement, beyond the minimum needed to produce it, the weaker that shift will be [1].
More pressure, less change. That is not a small adjustment to reinforcement theory. It is a reversal.
The idea got a name that describes it well: insufficient justification. The attitude moves precisely when the external reason for the behaviour is too small to cover it. Give people enough justification and nothing happens inside, because nothing needs to.
For a while the result looked like a clean knockout. Then a follow-up complicated it in a way that turned out to matter enormously.
In 1967, Linder, Cooper and Jones ran a version in which they varied not just the money but whether the person felt they had any real say in whether to take part [3]. When people felt they had chosen freely, the inverse pattern held: more money, less attitude change. When people felt they had no choice, the pattern flipped back to the ordinary direction. More money, more change.
Read that twice, because it relocates the whole theory. The active ingredient is not the money. It is the feeling of having chosen. The money only matters because of what it does to the story a person can tell about why they chose.
That finding is the reason the 2024 replication result, which we will get to, is so uncomfortable. The thing that failed to appear in 2024 was choice.
The Objection That Arrived Three Years Later and Never Left
Before the theory could be challenged on its own terms, it got a challenge from methodology, and the challenge has never been fully answered.
In 1962, Martin Orne published a paper about what he called demand characteristics [4]. His observation was simple and slightly devastating. A person who volunteers for an experiment is not a neutral measuring instrument. They are a person in a strange room trying to work out what is expected of them, and most of them would quite like to be helpful.
Orne's point applies to the peg study with uncomfortable directness. A man who has just been paid one dollar to say the task was fun, and who is then asked by an interviewer how much he enjoyed it, is in a socially awkward position. Saying it was miserable makes him someone who lies to strangers for a dollar. Saying it was fine makes the whole thing hang together.
That is not attitude change. That is a man not wanting to look bad.
Twenty-two years later, Rosenfeld, Giacalone and Tedeschi pressed the same objection as a formal alternative account, arguing that impression management could explain effort justification results without any internal discomfort at all [5]. Baumeister and Tice asked directly whether self-presentation, rather than an internal state, was doing the work in forced compliance studies [6].
The 1959 design does something about this. The interview was run by a different person, in a different room, framed as unrelated to the experiment. That is a real precaution and it was ahead of its time. Whether it is sufficient is a question the field is still arguing about, and you will see it resurface in the 2024 commentary.
Bem's Move: You Read Your Own Mind the Way Strangers Do
The most serious theoretical challenge arrived in 1967, and it did not attack the data at all. It accepted every number and offered a different reason for them.
Daryl Bem proposed self-perception theory [7]. His argument is that people do not have privileged access to their own attitudes in the way we assume. When you ask someone what they think about something, they often do not consult an inner record. They look at their own recent behaviour and infer what they must think, using exactly the evidence an outside observer would use.
Apply that to the peg study and the discomfort disappears from the explanation.
A man who told a stranger the task was fun, for one dollar, looks at his own behaviour and reasons the way you would about a stranger. He said it was fun. He was barely paid. People do not say things for a dollar unless they half mean them. So he probably enjoyed it a bit.
A man paid twenty dollars runs the same inference and gets a different answer. He said it was fun. He was paid well. That explains it. No conclusion about his attitude needed.
Same outputs. No dissonance, no discomfort, no drive. Just an ordinary inference running on visible evidence.
To show this was not a story he had made up after the fact, Bem ran the inference in people who had no stake at all. He described the experiment to observers who had never done the task and asked them to guess how the participant felt. The observers produced the same pattern the real participants had produced. Somebody told about a man paid one dollar guessed that man had enjoyed it more than a man paid twenty.
That is a strong result and it should be uncomfortable for dissonance theory. If a person who felt no discomfort whatsoever can reproduce the finding from the outside, the discomfort may not be the thing generating it.
Bem's paper is, by a wide margin, the most cited item in this entire literature. It did not go away.
There is a family resemblance here to what the split-brain experiments later showed in a much more literal form: a person confidently explaining their own behaviour with a reason that could not possibly be the real one, and having no sense that they are doing it.
The Experiments That Were Supposed to Settle It
Through the late 1960s and into the 1970s the field tried to arbitrate. The idea was to find a situation where dissonance theory and self-perception theory made different predictions, then look.
Bem and McConnell argued the case from the salience of a person's attitude before the manipulation [8]. Snyder and Ebbesen built a test around whether people were made aware of the dissonance itself [9]. Studies accumulated. Both camps claimed them.
In 1975, Anthony Greenwald wrote the paper that should have ended the arms race, with a title that gives away the conclusion: on the inconclusiveness of crucial cognitive tests of dissonance versus self-perception theories [10].
His argument was that the two theories, as stated, predicted the same observable outputs in essentially every design anyone had built. Running more of those designs could not separate them. What was needed was a measurement of something other than the final attitude.
That turned out to be exactly right, and it is where the story gets interesting.

The Pill
If dissonance is a felt state and self-perception is a cool inference, then the felt state should behave like a felt state. It should be possible to misattribute it.
In 1974, Mark Zanna and Joel Cooper built a study around that [11]. Everyone swallowed a placebo capsule before doing a standard induced compliance task, and the only thing that differed between groups was what they had been told the capsule would do. Tense, said one version. Relaxed, said the second. Nothing at all, said the third.
The prediction is precise and slightly strange, which is what makes it a good test.
If you have been told the pill makes people tense, then the unpleasant feeling you notice after writing something you disagree with has an available explanation that has nothing to do with you. It is the pill. So you should not need to change your attitude, and you should not.
If you have been told the pill relaxes you, then feeling tense anyway means the discomfort must be strong, since it got through a relaxant. So you should change more.
That is what happened. Participants told to expect tension showed no attitude change. Participants told to expect relaxation showed more than the control condition.
Self-perception theory has no obvious reason to predict this. An inference from your own behaviour should not care what you were told about a capsule. Something that behaves like arousal, and can be redirected onto a pill, is doing work.
A related result four years later found that dissonance arousal could be transferred onto an unrelated stimulus, behaving like a general undifferentiated state rather than a specific feeling about the belief in question [12].
Bem was not refuted. He was outflanked. The field found a way to measure something other than the final attitude, exactly as Greenwald had said it would have to.
Sweat, Discomfort and a Drink
Once the arousal question was open, people went looking for it directly.
In 1983, Croyle and Cooper measured skin conductance during an induced compliance procedure and found elevated arousal in the condition where dissonance theory said it should be [13]. Note that study, because it is the one that gets replicated in 2024 and it is not the peg study.
Elkin and Leippe found a similar arousal link, along with an effect they described as a reluctance to be reminded of the discrepancy [14]. Losch and Cacioppo found the sympathetic arousal too, but argued for a subtly different reading: attitudes changed to reduce negative feeling rather than to reduce arousal as such [15].
That difference matters more than it sounds. Arousal is a body state. A negative feeling is an interpretation of one, and only the second is the kind of thing a person can argue themselves out of.
In 1994 Elliot and Devine took it seriously enough to measure the discomfort directly rather than inferring it, and found that people in dissonance conditions reported feeling worse, that the discomfort preceded the attitude change, and that changing the attitude reduced the discomfort [16]. Harmon-Jones later showed the negative affect appeared even when the behaviour produced no aversive consequences for anyone, which had been proposed as a necessary condition [17].
The most memorable entry in this line is also the simplest. In 1981, Steele, Southwick and Critchlow gave people alcohol after a dissonance manipulation. The attitude change went away [18]. If the discomfort is what drives the repair, and you switch off the discomfort, the repair should not be needed. It was not.
The measuring did not stop, and it has not all gone dissonance theory's way. A preregistered study in 2021 went after both halves at once, using skin conductance for the arousal and heart rate variability for the reduction. It found weak evidence for the arousal and none for the reduction [19]. That is a result about the instruments as much as the theory, but it is not one anybody wanted. A different approach, pointing the other way, leaves the laboratory altogether. People who score high on dissonance arousal turn out to get less exercise, and people who score high on dissonance reduction get more, measured in one study across a year of GPS-tracked cycling [20].
None of this proves Festinger's specific account. It does establish that something is being felt, that it is unpleasant, and that it can be manipulated by things that have nothing to do with the belief in question.
Four Ways Out, and the One You Have Never Heard Of
Popular accounts of dissonance almost always describe two exits. Change what you think, or change what you do.
The literature describes more, and one of them turns out to matter for reasons that go well beyond tidiness.
The fourth route is trivialization, described by Simon, Greenberg and Brehm in 1995 as the forgotten mode of dissonance reduction [21]. Instead of changing the belief or the behaviour, you shrink the importance of the whole thing. It was one dollar. It was one afternoon. It did not really matter.
You have almost certainly done this. It is the cheapest exit available, because it costs you no belief and no behaviour.
It also creates a measurement problem that runs through the entire field. If a study only measures attitude change, and a participant reduces their discomfort by deciding the issue is unimportant, the study will record that participant as showing no effect. The dissonance was there. The resolution happened. The instrument was pointed somewhere else.
There are other routes still. People add consonant thoughts, bringing in new considerations that make the behaviour make sense. People also affirm themselves somewhere else entirely, shoring up a different corner of the self-image and leaving the actual contradiction untouched. That one grew into a research tradition of its own, with instruments for measuring how readily a given person reaches for it [22]. And people practise what one recent line of work calls deliberate ignorance: avoiding the information that would create the conflict in the first place [23].
A review of dissonance reduction strategies gives a sense of how many named routes there now are [24]. The important point for reading any individual study is that a person with several exits available will take the cheap one, and the cheap one is often invisible to the measure.
The Family of Experiments
The peg study is one member of a family, and the members have not aged identically. Arguments about dissonance slide between them constantly, usually without announcing it.
Three years before the pegs, Jack Brehm published the free choice paradigm [25]. Ask people to rate a set of objects, offer them a choice between two they rated about equally, and re-rate afterwards. The chosen item goes up. The rejected item goes down. The choice appears to have created a preference gap that did not exist before it.
The same year as the pegs, Elliot Aronson and Judson Mills published the severe initiation study [26]. Women who had to complete an embarrassing procedure to join a discussion group afterwards rated the group, which was deliberately made dull, more favourably than women admitted easily. Effort justification: if it cost you a lot, it must have been worth something.
In 1968, Knox and Inkster took the idea to a racetrack [27]. They asked bettors how confident they were about their horse, some just before placing the bet and some just after. The people who had already handed over the money were more confident. Same horse, same race, same information. The only thing that changed was that the decision had become irreversible.
They all share a shape. In every one the person has already done something they cannot take back, and in every one the belief moves afterwards to make the act look sensible. What differs is only the irreversible thing: a lie, an initiation, a bet, a choice.
Aronson spent decades arguing that the theory had been stated too broadly [28], and he was still arguing it late in his career [29]. On his account dissonance does not bite whenever two thoughts clash. It bites when the behaviour threatens something you believe about yourself, which is why lying about the pegs stings and being wrong about them does not.
Axsom and Cooper then took effort justification out of the laboratory. People made to work hard for a weight loss programme lost more weight than people given an easier version of the same programme [30]. The effort itself was doing therapeutic work, which is an uncomfortable thing to learn about your own sense of what treatment is.
The effort question is still live, and messy, because two literatures say opposite things. Effort justification says working hard for something makes you value it more. Effort discounting says it makes you value it less. A 2024 study reconciled them by adding perceived control. When people felt in control, more effort raised the reward's value. When they did not, more effort lowered it [31]. Control again, exactly as in 1967.
A 2025 series then found that effort does reliably make a task feel more meaningful, up to a point. What it did not find was any evidence that dissonance was the reason [32]. The effect looks real and the mechanism looks like something else, which is a sentence that could be written about a lot of this field.
Dissonance even runs second-hand. Watching somebody from your own group act against an attitude you share produces it in you, and a preregistered meta-analysis now supports the effect [33]. That one is hard to explain as self-justification, because there is no act of yours to justify.
Then there is hypocrisy induction, which is the most practically useful member of the family. Fried and Aronson had people publicly advocate a behaviour and then made them think about their own failures to do it [34]. The gap between the sermon and the record is uncomfortable, and people close it by changing the record. Stone and Fernandez reviewed the technique as a behaviour change tool [35].

The Confound That Nearly Took Fifty Years With It
In 2010, Keith Chen and Jane Risen published an argument about the free choice paradigm that is genuinely alarming if you follow it through [36].
Their point is statistical and it is very hard to escape.
Suppose you rate two items as equally attractive. That rating is not a perfect readout of your preference. It is a noisy measurement of it. Underneath the tie you reported, you almost certainly have some slight real preference, and when you are then forced to choose, that hidden preference is what breaks the tie.
So the choice does not create the gap. It reveals it. And when you re-rate afterwards, the gap you find was there all along, hiding under the measurement error of the first rating.
That is not a small caveat. Chen and Risen showed the artifact was large enough to produce the classic result with no attitude change whatsoever, and it applies to a paradigm that had been generating findings since 1956. It reached into some very well known results, including one that had been reported as evidence that the effect appears in preschool children and in capuchin monkeys [37].
To their credit, the authors of that study went and dealt with it. Egan, Bloom and Santos designed a version in which the subject made a choice without knowing what they were choosing between, a blind two-choice procedure in which no hidden preference could break the tie because there was nothing visible to prefer [38]. The effect survived.
Izuma and Murayama then wrote a methodological review setting out exactly which free-choice designs are vulnerable and which are not [39]. And in 2021, Enisman, Shpitzer and Kleiman pooled 43 studies covering 2,191 people, every one of them using a design built to be immune to the artifact. The effect was still there, at d = 0.40 with a confidence interval of 0.32 to 0.49, and they found no sign that publication bias accounted for it [40]. Choosing changes what you want. It does not merely reveal it.
That sequence is the field working properly. A serious objection was raised, it was taken seriously, designs were fixed, and the finding was retested rather than defended. Whether the change lasts is a separate question, and at least one study has looked at that specifically [41].
It is also the standard the next section has to be judged against.
Children, Monkeys and the Question of How Old This Is
The Egan study deserves a closer look, and so do its limits.
Preschool children and capuchin monkeys were given a choice between two things they had shown equal liking for: two stickers for the children, two colours of the same sweet for the monkeys. Afterwards, offered the rejected option against a fresh third option they had originally liked just as much, both groups went for the new one. They had devalued the thing they turned down [37].
These are small samples of children and of animals, and they should be read as such. A capuchin devaluing a sweet is not a person rewriting a belief about themselves. The authors were careful about this and described it as decision rationalization rather than as cognitive dissonance in the full human sense. A later review looked specifically at what dissonance reduction in nonhuman animals can and cannot tell us about the human theory [42].
What it does suggest, if the finding holds, is that whatever produces the effect does not require the elaborate self-concept that later versions of the theory lean on. Something simpler is available to a four-year-old and to a monkey.
Culture complicates the picture without undermining it. Hoshino-Browne and colleagues found that dissonance appears for people from Eastern and Western backgrounds, but for different kinds of choice: personal choices for one group, choices made on behalf of a friend for the other [43]. The mechanism looks general. What counts as a threatening inconsistency does not.
Dissonance also runs through groups rather than only individuals. Matz and Wood found that simply being in a group where others disagree with you produces discomfort and a push toward resolution, either by changing your view, changing theirs, or leaving [44].
Inside the Brain
By the late 2000s the tools existed to ask where in the brain this happens. The answers are more solid than the patient work that follows them and less specific than they first look.
In 2009, van Veen and colleagues put people through an induced compliance procedure in a scanner, having them say positive things about the experience of being in an MRI, which is genuinely unpleasant. Activity in the dorsal anterior cingulate cortex and the anterior insula during the counterattitudinal statements predicted how much the person's attitude subsequently moved [45].
Both regions are associated with conflict detection and with the felt quality of unpleasant states. That is roughly what a physical implementation of Festinger's idea would look like if you had to guess in advance.
The following year, Izuma and colleagues studied choice-induced preference change with a control condition designed to rule out the artifact Chen and Risen had described, and found changes in striatal activity tracking the preference shift [46]. Jarcho, Berkman and Lieberman caught rationalization in the act, scanning people while they were making the choice rather than after it [47]. Kitayama and colleagues examined choice justification and found the neural signature varied with cultural background in a way that echoed the behavioural work [48].
Correlation is not the end of the argument, and in 2015 Izuma and colleagues went further. Disrupting the posterior medial frontal cortex changed the size of the preference shift rather than merely accompanying it [49]. That is a causal claim about a specific region and it is one of the strongest pieces of evidence in the modern literature. A 2026 scoping review collects the imaging work, and the fair summary is that the regions implicated are largely the ones that light up whenever anything is in conflict. Suggestive rather than specific [50].
The Patients Who Forgot the Choice and Kept the Preference
The most surprising result in this whole area came from people whose memories do not work.
In 2001, Lieberman, Ochsner, Gilbert and Schacter ran the free choice procedure with amnesic patients. These are people who cannot form new explicit memories, so a short while after making a choice they do not remember making it.
They showed the preference change anyway [51].
Think about what that rules out. A person who has no memory of choosing cannot be justifying a choice they know they made. They cannot be managing an impression of a decision they cannot recall. Whatever changed the preference did not require conscious access to the event that caused it.
Except the picture is not that clean.
In 2017, Chammat and colleagues reported that dissonance resolution depends on episodic memory, working with patients whose memory systems were impaired [52]. In 2021, Tandetnik and colleagues found that resolution depends on executive function and on frontal lobe integrity [53].
Those three results do not sit comfortably together. The first says the repair runs fine with no explicit memory of what triggered it. The other two say it leans on exactly the memory and frontal machinery the first one managed without. They use different patient groups, different tasks and different measures, and reconciling them is an open problem rather than a solved one.
That has consequences for how confidently anyone can describe the mechanism. We have converging imaging evidence about where conflict is registered, one causal disruption result, and a set of patient findings that disagree. That is a live research area, not a settled account, and the popular framing of dissonance as a well-understood brain process is running ahead of what is actually known. The same overconfidence shows up in how people describe their own reasoning generally, which is the territory of the Dunning-Kruger effect and its own history of being oversold.
Sixty-Seven Years, in Order
Before the modern part of the story, it helps to see the shape of it.
The rhythm of a healthy scientific argument runs down that list. A claim, an alternative, a measurement that separates them, a methodological attack, a fix, a retest. Then, near the bottom, something that has not been resolved yet.

2024: Thirty-Nine Labs and Four Thousand Eight Hundred and Ninety-Eight People
In 2024, David Vaidis and a very large group of collaborators published the biggest test the induced compliance paradigm has ever had [2].
The numbers first. Thirty-nine laboratories. Nineteen countries. Four thousand eight hundred and ninety-eight participants. Preregistered.
The design was a constructive replication of Croyle and Cooper's 1983 experiment, the skin conductance study from earlier in this article. That detail is the one most likely to be misreported, so here it is flatly: they did not replicate the one dollar peg study. They replicated a different, later member of the same family.
Participants wrote a counterattitudinal essay. Some were given the standard high-choice framing, in which the experimenter makes clear the decision is theirs. Others were given a low-choice framing. If the induced compliance account is right, the high-choice group should shift their attitudes more, because only they have a decision to justify.
The choice effect was not small. It was absent. A d of minus 0.03 with an interval running from minus 0.10 to 0.04 is not an underpowered study failing to reach significance. It is a large study placing a tight bracket around zero.
The second row is the one that complicates it. Writing a counterattitudinal essay did move attitudes, by a modest but clearly non-zero amount, compared with writing a neutral one. Something happened. It just did not depend on whether people felt they had chosen.
The authors' own conclusion is measured. Their results, they write, call into question whether the induced compliance paradigm provides evidence for cognitive dissonance that can be relied on, and suggest that choice may not be necessary for attitude change in this setting.
That is awkward for a precise reason. Since Linder, Cooper and Jones in 1967, perceived choice has been the ingredient the theory hangs on. It is why one dollar beats twenty. Remove choice and the standard account has nothing left to explain the direction of the original effect.
What a Null Actually Establishes
Here is where a lot of coverage of replication failures goes wrong, and it goes wrong in both directions.
The published response was immediate and substantive. Advances in Methods and Practices in Psychological Science ran commentaries alongside the paper, which is what good journals do with a result of this weight.
Eddie Harmon-Jones and Cindy Harmon-Jones, who have spent decades working on dissonance, argued that the replication's implementation did not create dissonance in the first place [54]. Their case is about the details of the manipulation: whether the essay topic mattered enough to the participants, whether the choice framing was convincing, whether the procedure produced the conditions the theory specifies. If the manipulation did not induce the state, then a null result is not evidence about the state.
That argument is easy to dismiss as special pleading and it should not be. It is testable in principle, and it is precisely the argument the field accepted when Egan and colleagues answered Chen and Risen by building a better design. The right response is another study, not a verdict.
Wilson Cyrus-Lai, Warren Tierney and Eric Uhlmann took on the general question instead: what can anyone conclude when a classic finding fails to replicate [55]. Their answer is uncomfortable for everyone. A failed replication is real information and it is weaker information than either camp wants it to be, because the space of things that could have gone wrong is large and includes the original.
So what should you take from 2024?
Something narrow and specific. The standard laboratory recipe for producing dissonance, run at enormous scale with modern methods, did not produce the effect its central variable predicts. That is a serious problem for the paradigm. It is not, by itself, a verdict on whether people rationalise their choices, because the paradigm is one way of asking and the answer arrives from other directions too: the arousal work, the causal brain disruption, the artifact-free meta-analysis, and the applied trials in the next section.
What it does mean is that any confident sentence about dissonance that rests only on induced compliance is now standing on ground that moved. That includes a lot of textbook writing. Recent overviews have started to reckon with this [56], and general reviews of attitude change have long noted the paradigm's fragility [57].
The Theories That Grew Out of the Argument
One consequence of sixty-seven years of pressure is that the original theory has been rebuilt several times, and the versions differ in what they say the discomfort is about.
Stone and Cooper proposed the self-standards model, in which what matters is the standard you judge your behaviour against, and the discomfort depends on which standard is available at the moment [58]. They later showed that whether self-esteem protects you or exposes you depends on how relevant the threatened attribute is to your sense of yourself [59]. Behind both sits a broader claim that people run several self-esteem repair mechanisms at once and trade between them, so blocking one route sends the work down another [60].
Harmon-Jones and colleagues developed the action-based model, which asks what the discomfort is for [61]. On this view conflicting thoughts are a practical problem before they are an emotional one, because somebody holding both cannot act cleanly on either. Resolving it is not vanity. It is how you get unstuck. That framing also predicts things the other models do not, and the exercise finding above is one place it has been tested. Harmon-Jones and Harmon-Jones summarised where fifty years of development had left the theory [62].
Those two are not variations on a theme. The self-standards model says dissonance is fundamentally about the self. The action-based model says it is about getting things done. They point at different experiments, and the field has not chosen between them.
Shultz and Lepper went around both of them, modelling dissonance reduction as constraint satisfaction in a neural network. A system doing nothing but settling into its most consistent available state reproduced a good deal of the classic findings [63]. No discomfort, no drive, no self. If a network with none of those can produce the behaviour, the behaviour is not strong evidence that people have them.
Three recent proposals make genuinely different bets and deserve separating from the pile. Social verification theory moves the problem outward, arguing that what people cannot tolerate is inconsistency with those around them rather than inside their own heads, and that chronic social inconsistency behaves rather like chronic rejection [64]. A 2025 Psychological Review paper goes the other way and adds a switch: people rationalise only when they accept the outcome in the first place. That would explain why the effect keeps appearing and vanishing across studies that look identical on paper [65]. The third asks about the threshold instead: how much a conflict has to matter before somebody will spend real effort resolving it rather than shrugging [66]. If most conflicts fall below that line, a good deal of laboratory silence would follow.
Six models for one phenomenon is a lot, and the number is itself a finding. A theory that had been pinned down would not need this many versions of itself.
There is also a persistent finding that dissonance moves explicit attitudes while leaving implicit ones untouched [67], and that a gap between the two produces its own kind of discomfort and its own extra processing [68]. That suggests the repair may be more superficial than the theory's grander readings imply. You change what you say you think. Something underneath may not move at all.
Where the Idea Still Earns Its Keep
If the paradigm is fragile, why does anyone still use the theory?
Because the applied version has an evidence base that does not depend on the laboratory recipe that failed.
The strongest case is eating disorder prevention. The approach is hypocrisy induction with a specific target: participants argue out loud against a thin ideal they have partly internalised, in their own words, in front of other people.
The first randomised trial, in 2001, put 87 young women with body image concerns through it. They came out with lower thin-ideal internalisation, less body dissatisfaction, less dieting and fewer bulimic symptoms, and the gains were still there four weeks later [69]. The authors also reported something they had not expected. The control group, given ordinary weight management advice, improved on some measures too.
The programme kept being run, and eventually there were enough trials to pool. The 2019 meta-analysis put the effect on eating disorder symptoms at around d = 0.31, with smaller effects on negative affect, and then went looking for what separates a programme that works from one that does not [70]. More dissonance-inducing activities helped. So did running the sessions in person rather than online, in larger groups, led by trained clinicians rather than by researchers.
Then one of the moderators came out backwards. The effects were larger when participants were paid to take part.
Read that against the peg study. Payment is supposed to supply a reason and drain the dissonance away, which is exactly what twenty dollars did in 1959. Here it made the intervention work better. The authors flag it as unexpected and leave it there, because nobody has an explanation.
The 2021 meta-analysis asked the harder question. Not whether questionnaire scores move, but whether fewer people go on to develop an eating disorder. The answer was yes, at an odds ratio of 1.64 with a confidence interval of 1.09 to 2.46, which the authors translate into a 54 to 77 percent reduction in future onset [71]. That figure comes from the applied literature rather than the laboratory one, which is the opposite of where a reader would expect the best evidence to be.
Pointed at other targets the record is patchier. A 2025 randomised trial gave heavy-drinking college students a counter-attitudinal advocacy task. It did not reduce how much they drank. It did reduce the number of alcohol-related problems they reported, which the authors describe as a harm reduction effect on consequences but not on consumption [72].
A second 2025 trial built a dissonance-based intervention against weight stigma and randomised 325 students to it. Stigma scores fell in every condition, including the control. The dissonance arms were not reliably better, and the authors report limited evidence that dissonance was induced at all [73].
That is the 2024 problem in miniature: a procedure built straight from the theory, run properly, that could not show it had produced the state it was designed to produce.
Then there is the everyday case that most people will recognise, which researchers call the meat paradox. Most people say they care about animals. Most people also eat them. That is a textbook dissonant pair, and the reduction strategies people use are exactly the ones the theory predicts: changing behaviour, changing the belief, or shrinking the importance of the conflict. One study mapped the strategies across dietary groups and found them used most heavily by the people with most to justify. Omnivores leaned hardest on denial of animal suffering and on a sharp mental split between animals you eat and animals you do not, and the use of both thinned out toward vegans [74].
Another tested which arguments actually generate the discomfort. Animal rights and environmental messages raised dissonance and shifted attitudes toward eating fewer animal products. The health argument did neither [75]. Telling people it is bad for them does not create a moral conflict, because there is no moral conflict in it.
The strangest of the three ran 1,501 people and found that being angry at somebody else works as a dissonance reducer. Meat eaters reminded of factory farming and then given a third party to blame reported less guilt and rated their own moral character higher. Self-affirmation removed the effect, which is what you would expect if the outrage was doing repair work on the self [76].
There is a result here for anyone who teaches. A 2025 study had people practise retrieving the content of persuasive texts. The testing improved how much they learned and left their attitudes exactly where they were [77]. Understanding a message and being moved by it are separate processes, and only one responds to studying harder.
A 2025 review sets out why correcting a belief so often fails to change behaviour. A corrected belief only moves what people do when three things hold: it is tied to something the person actually wants, the inferential path to action is short, and the link survives in memory to the moment of acting [78]. Most corrections fail at least one.
That problem is a close cousin of the one this whole article keeps circling. Beliefs are not held in isolation and they are not updated on evidence alone. The mechanism that protects a belief before it is challenged is described in a separate piece on confirmation bias. What dissonance describes is the repair crew that arrives afterwards, once the damage has already been done by your own behaviour.
The Machine That Changed Its Mind
The strangest entry in the modern literature was published in 2025, and it is extremely easy to overinterpret.
Steven Lehr and colleagues ran two preregistered studies on GPT-4o [79]. The model was asked to write an essay about a political figure, either positive or negative. Afterwards its stated attitude toward that figure had shifted in the direction of the essay it had just written.
That alone is interesting but explicable. A model conditioned on text it just produced might well produce consistent text afterwards.
The second finding is the odd one. When the model was offered an illusion of choice about which essay to write, the shift was sharply larger. The authors describe this as a functional analog of humanlike selfhood and say plainly that the mechanism is unknown. They defended the interpretation against critics in a published reply [80].
Held next to the 2024 result, that pair is uncomfortable. In four thousand eight hundred and ninety-eight humans, the choice manipulation did nothing. In a language model, the same manipulation sharply increased the shift.
Resist the two easy readings. This is not evidence that the model has an inner life, and the authors do not claim it does. It is also not evidence that the human effect is real after all, because a system trained on human text reproducing a pattern found in human text is not independent confirmation of anything about humans.
What it does raise is a genuinely strange possibility: that the pattern is in some sense a property of how self-descriptions get generated from prior behaviour, which is much closer to Bem's 1967 argument than to Festinger's. Bem said people infer their attitudes from their own behaviour the way an observer would. A language model with no discomfort, no arousal and no self, doing exactly that, is not the result dissonance theory would have ordered.
Nearly sixty years on, the argument between Festinger and Bem has acquired a participant neither of them anticipated.
What the Peg Study Actually Left Us
So where does this leave an experiment that fits on two pages and involved sixty men in a room in 1959?
Some things are solid. The numbers in that table are the numbers. The direction of the effect was the opposite of what the reigning theory predicted, and getting that right in advance is the strongest thing a theory can do. People do rationalise decisions they cannot undo, in many paradigms, across many decades, and the versions of those paradigms that were rebuilt to remove known artifacts still show it.
Some things are genuinely open. Whether the standard laboratory procedure reliably produces the state it claims to produce is now an open question with a very large data point against it. Whether the discomfort is arousal or negative affect or something else is unsettled. Whether the brain evidence tells a single story is unsettled, and the patient findings actively disagree. Whether Bem's account was ever really defeated, or just outmanoeuvred on measurement, is a fair question to reopen in 2026.
And some things that are said constantly are simply not true. The peg study did not fail to replicate, because it has not been directly repeated at scale. The theory has not been debunked, and nobody involved in the 2024 replication says it has. The experiment did not show that people will believe anything for money, it showed the reverse. It did not measure enjoyment, it measured four things and found one.
The reason to care about all of this is not really about pegs.
It is that a very large amount of what you believe about yourself was assembled the same way. You did something. It was too late to undo. And a quiet process you had no access to went to work on the belief instead, and handed you back an opinion that felt like it had been there all along.
That process does not feel like reasoning while it is happening. It feels like noticing. Which is exactly what makes it hard to catch in yourself, and exactly why it took sixty men, forty-eight pegs and sixty-seven years of arguing to get this far. The same machinery that quietly rewrites an attitude also quietly rewrites what you remember, which is the subject of separate pieces on how false memories form and on memory as reconstruction.
One last thing, and it is about the science rather than the psychology.
A field that publishes a study, watches it become a classic, then spends decades attacking it, funding thirty-nine labs to test it properly, and printing the objections alongside the result, is not a field in crisis. That is what the process is supposed to look like. The 1959 result was small, and it was interesting, and it was taken far too seriously for far too long by people who had never read the table. Now it is being read properly. Whatever survives that will be worth knowing.
Frequently Asked Questions
What is cognitive dissonance in simple terms?
It is the discomfort you feel when two things you hold at once do not fit together, usually a belief about yourself and something you have just done. Leon Festinger proposed in 1957 that this discomfort works like a drive: it is unpleasant enough that you are pushed to get rid of it. Since you often cannot undo the behaviour, the cheaper repair is to adjust the belief. The important and counterintuitive part is that this does not feel like changing your mind. It feels like realising something. That is why it is so hard to notice in yourself and so easy to spot in other people.
What did the Festinger and Carlsmith experiment actually find?
Sixty male Stanford undergraduates spent an hour on a deliberately boring task, thirty minutes moving spools and thirty minutes turning forty-eight pegs a quarter turn at a time. Some were then paid one dollar or twenty dollars to tell the next participant it had been enjoyable. On an eleven-point enjoyment scale the men paid one dollar rated the task at plus 1.35, the men paid twenty at minus 0.05, and a control group who never lied at minus 0.45. Three other questions were asked, about learning, scientific importance and willingness to repeat, and those produced much smaller differences. The whole famous result is one row of a four-row table, from twenty men per condition.
Why did one dollar change more minds than twenty?
Because twenty dollars was a good reason and one dollar was not. In 1959 money, twenty dollars was worth something in the region of two hundred dollars today, which is more than enough to explain why a student would say something he did not mean. The man paid one dollar has no such explanation available, so he is stuck holding two thoughts that do not fit, and the one that can move is his opinion of the task. This is called insufficient justification. A follow-up in 1967 showed the pattern only appears when people feel the decision to speak was genuinely theirs, which means the real active ingredient is choice rather than money.
Did the cognitive dissonance experiment replicate?
Not directly, and this is widely misreported. In 2024, thirty-nine labs across nineteen countries ran 4,898 participants through a large preregistered replication, but of Croyle and Cooper's 1983 study rather than the 1959 peg experiment. The choice manipulation, which the theory says is essential, produced nothing: a Cohen's d of minus 0.03 with a confidence interval from minus 0.10 to 0.04. Writing a counterattitudinal essay did still shift attitudes compared with a neutral one. Researchers who work on dissonance published a commentary arguing the procedure never induced dissonance in the first place, and another commentary examined what a failed replication of a classic can establish at all. The argument is live and unresolved.
Is cognitive dissonance the same thing as being a hypocrite?
No, and the difference is about awareness. Hypocrisy is a description other people apply to a gap between what you say and what you do. Dissonance is the internal discomfort that gap creates, and the repair usually happens without your noticing, which is why it does not feel like hypocrisy from the inside. There is a research technique called hypocrisy induction that deliberately makes the gap visible, and it is used in behaviour change programmes precisely because people who are made to see their own inconsistency tend to close it by changing the behaviour rather than the words.




