Introduction

Here are three numbers. 2, 4, 6.

They follow a rule. I know the rule and you do not. You can test as many other triples as you like, and each time I will tell you whether your triple follows the rule or not. When you are sure you have it, say the rule out loud.

Most people look at 2, 4, 6 and think: ascending by twos. So they try 8, 10, 12. Yes, that follows the rule. They try 20, 22, 24. Yes. They try 100, 102, 104. Yes again. Three for three. They announce the rule with real confidence.

The rule was any three ascending numbers.

You could have found it in one move. All you had to do was try something you expected to fail. 1, 2, 3 would have done it. So would 5, 40, 900. Instead you spent your questions collecting yeses, and every yes felt like progress while telling you almost nothing.

This is Wason's rule discovery task [1], run in 1960, and it is where nearly every article about confirmation bias begins. The standard reading is that it proves something bleak about people. We do not test our beliefs. We audition evidence for them.

That reading has been wrong since 1987, and the correction is the reason this article exists.

Glowing amber spheres in deep indigo space with faint lattice background.

What Confirmation Bias Actually Names

Start with a definition that will survive the rest of the article, because several popular ones will not.

Confirmation bias is the tendency to look for, weigh, and remember information in a way that fits what you already think. Not deliberately. Not because you are lying to yourself. It happens upstream of anything you would recognise as a decision.

The phrase itself arrived later than the phenomenon. Mynatt and colleagues put it into print in 1977, in a study that built a small artificial universe of physics for people to investigate [2]. Participants could fire particles at objects and watch what happened. They mostly ran the experiments that would confirm what they already suspected, and when a result contradicted them, they often kept the theory and doubted the result.

Nickerson's 1998 review is still the reference point for the whole area, and it did something the popular versions rarely do [3]. It took the phenomenon seriously as a family of related behaviours rather than a single trick of the mind, and it was careful about which of them are actually irrational.

That care matters more than it sounds. Because the word bias smuggles in a verdict.

A bias, in ordinary speech, is a flaw. Something that pushes you off true. The trouble is that a great deal of what gets filed under confirmation bias is not a flaw at all, and the clearest demonstration of that is the very experiment we started with.

1987: The Year the Famous Result Stopped Meaning What It Said

Go back to 2, 4, 6.

You guessed ascending by twos. You tested 8, 10, 12. Why is that a bad move?

The usual answer is that you were seeking confirmation instead of trying to falsify. Popper said scientists should try to break their own theories, you did the opposite, and therefore you reason badly.

Klayman and Ha showed this answer does not hold up [4]. Their argument is formal and it is worth following, because it changes what the whole field is about.

Testing a case you expect to be positive is called a positive test strategy. It is not the same thing as seeking confirmation, although in Wason's task the two happen to coincide. Klayman and Ha asked a different question: across the range of situations people actually face, how well does the positive test strategy perform?

The answer is that it usually performs very well. Suppose you are trying to work out what causes your headaches, or which of your students is struggling, or what makes a bridge fail. In cases like these the thing you are hunting is rare. It covers a small slice of everything that could happen. When the target is small, testing cases inside your current hypothesis is the fastest way to find its edges. Disconfirmation arrives on its own, quickly, because a narrow hypothesis in a wide world keeps bumping into things it did not predict.

The strategy only fails in one specific arrangement. It fails when your hypothesis sits entirely inside the true rule.

Look at what Wason built. His rule was any three ascending numbers, which is enormous. Your guess, ascending by twos, fits completely inside it. Every triple that satisfies your guess also satisfies his rule. A positive test literally cannot fail. The task was constructed so that the strategy which normally works could produce nothing but yeses.

That is not a demonstration of irrationality. It is a demonstration of a sensible heuristic dropped into the one environment built to defeat it.

Oaksford and Chater made a parallel argument about Wason's other famous task, the four card selection problem [5]. People who fail it are not violating logic at random. They are picking the cards that would be most informative if the world were arranged the way worlds usually are.

The pattern in both cases is the same. Take a strategy people use, put it in a setting where it happens to fail, and you can make almost any reasoner look broken. The interesting question is never whether a strategy can be defeated. It is how often the world is arranged to defeat it.

None of this means confirmation bias is imaginary. Klayman himself spent a later paper carefully separating the varieties, some of which are genuinely costly [6]. The point is narrower and sharper. The founding experiment is not the proof it is presented as, and an article that opens with 2, 4, 6 and moves straight to "people are irrational" has skipped forty years of argument.

Hahn and Harris pushed the question further and asked what it would even mean to call a reasoning pattern biased [7]. To call something a deviation you need a standard to deviate from, and choosing that standard is a substantive claim, not a neutral setup. Peters goes further and asks what the tendency is for, on the grounds that something this consistent across people is more likely to be doing a job we have named badly than doing no job at all [8].

Hold onto that. We will come back to it with better evidence than philosophy.

1620
Bacon describes the tendency in Novum Organum
1960
Wason runs the 2 4 6 rule discovery task
1968
Wason introduces the four card selection task
1977
Mynatt and colleagues put the term into print
1979
Lord Ross and Lepper demonstrate biased assimilation
1987
Klayman and Ha reframe it as positive testing
1998
Nickerson publishes the landmark review
2019
Wood and Porter find the backfire effect elusive
2024
Pilgrim derives the bias from bounded Bayesian updating
2025
Leung and Urminsky name the narrow search effect

The history matters because the popular account froze in about 1979 and never updated.

Three Different Failures Wearing One Name

Ask most people what confirmation bias is and you get one thing. The research describes three, and they happen at different moments, in different systems, with different fixes.

You search selectively. That happens before you have seen anything.

You interpret selectively. That happens while you are reading.

You remember selectively. That happens days later, when the evidence is gone and only your reconstruction of it is left.

These are not three flavours of one process. They are three separate opportunities for the same preference to leave a mark, and a technique that closes one of them can leave the other two completely untouched. That is a large part of why the advice you have read does not work.

StageBiased searchBiased interpretationBiased memory
When it happensBefore you see evidenceWhile you read itDays later at recall
What goes wrongYou generate only situations that could agreeYou hold contrary evidence to a stricter standardYou retrieve the episodes that fit
Founding studyWason 1960Lord Ross and Lepper 1979Snyder and Cantor 1979
Scale of the evidenceMeta analysed by Hart 2009Shown with fabricated studies matched for qualityExtends to stereotype maintenance
What a fix must changeHow the question is framedThe standard applied to each sideThe filtering has already happened by retrieval
Does knowing about it helpBarelyBarelyNot tested directly

Take them one at a time.

Pale light filtered through three colored translucent layers in a dark chamber.

One: The Question You Ask Decides the Answer You Get

Before anyone shows you evidence, you have already narrowed what evidence you will see. You did it when you framed the question.

This is the oldest strand of the research and it goes by the name selective exposure. Frey collected the early work in 1986 and the pattern was already clear [9]. Given a choice of things to read, people lean toward material that agrees with a position they have already taken, and the lean gets stronger when they have committed publicly or cannot easily reverse the decision.

For a long time the effect had a reputation for being unreliable, appearing in some studies and vanishing in others. Hart and colleagues settled the question with a meta analysis in 2009 [10]. Pooling the literature, people were meaningfully more likely to choose congenial information than uncongenial information. The effect is real. It is also moderated by what people are trying to do: when accuracy genuinely matters to them, the pull weakens.

That moderation is easy to skip past and it should not be. It says the pull is not fixed. When there is a real cost to being wrong, and the person knows it, they read the other side more willingly. Which means selective exposure is partly about how much the answer actually has to be right.

The mechanics are visible in how a search unfolds. Jonas and colleagues gave people information one piece at a time and watched the bias compound as the sequence went on [11]. Each confirming piece made the next choice a little more slanted.

You might expect a group to fix this. Put several people around a table, let them argue, and the individual leanings should cancel.

They do not cancel. Schulz-Hardt and colleagues found groups searched more selectively than individuals, not less [12]. A group that already leans one way treats agreement as evidence, and a room full of agreement is very comfortable to sit in. This is one of the least intuitive findings in the whole area and it undercuts a lot of ordinary management advice.

The bias reaches further down than deliberate choice, too. Rajsic and colleagues found it in visual search, in where the eyes go when you are hunting a target on a screen [13]. Told to look for a red letter, people checked red things first even when checking colour was a worse strategy than checking position. That is not a belief being defended. It is one hypothesis being made cheap to test while every alternative is made expensive, which is what attention does before memory gets involved at all.

Think about what that implies for advice. If part of this operates in where your eyes go on a screen, then telling somebody to be more open minded is aimed at a level of the system that was never the problem. You cannot instruct a saccade.

Trope and Bassok had shown something adjacent back in 1982 [14]. When people gather information about another person, they favour questions whose answers would fit the hypothesis they were handed, over questions that would actually discriminate between two possibilities.

Notice what all of this has in common. Nobody in these studies is refusing to look at contrary evidence. They are simply not generating the situations in which contrary evidence could arrive.

Yes

No

Existing belief

Question framed

Narrow evidence returned

Does it agree?

Belief strengthened

Source judged weak

That loop is the shape of the whole problem, and it closes before anything you would call reasoning has started.

Two: The Same Evidence Read Two Ways

Now suppose the evidence arrives anyway. Somebody hands it to you.

In 1979 Lord, Ross and Lepper ran the study that defined this stage [15]. They recruited people with strong opinions about capital punishment, both for and against, and gave every one of them the same two pieces of research. One study supported deterrence. One study undercut it. The studies were fabricated for the experiment and matched for quality.

Everybody found the study that agreed with them convincing and the study that disagreed with them methodologically weak. That part is expected. What was not expected is what happened to their opinions. After reading a balanced pair of studies, people on both sides came away more confident than they had been before.

The same evidence pushed two groups further apart. That result named biased assimilation, and it is the reason "just show them the data" so often fails.

Ditto and Lopez put a name to the mechanism in 1992 [16]. They called it motivated skepticism, and the finding is elegant. People do not simply reject unwelcome evidence. They apply a stricter standard to it. Welcome evidence gets asked "can I believe this?" and unwelcome evidence gets asked "must I believe this?" Both questions feel like honest evaluation from the inside. Only one of them is easy to pass.

That asymmetry is worth watching for in yourself, because it has a distinctive feel. Scrutiny that arrives only when you dislike a result is not scrutiny. It is a filter with a good reputation.

Taber and Lodge extended this to political reasoning and found the same asymmetry, with an additional twist [17]. People spent more effort arguing against contrary evidence than absorbing supportive evidence, which means the extra work was going into defence rather than understanding.

And beliefs are stubborn even after their foundation is explicitly removed. Ross, Lepper and Hubbard gave people false feedback about their performance on a task, then told them plainly that the feedback had been randomly assigned and meant nothing [18]. The belief survived the debriefing. Once you have built an explanation for why you are good or bad at something, taking away the original evidence does not take away the explanation.

That is belief perseverance, and it is why a retraction so rarely undoes what the original claim did.

Three: The Memory That Edits Itself

The third stage is the quietest and, for anyone interested in how memory works, the most interesting.

Snyder and Cantor ran the classic version in 1979 [19]. People read a long description of a woman named Jane whose week contained a mixture of outgoing and reserved behaviour. Two days later they came back and were asked to assess her suitability for a job. Half were asked about a job needing an extrovert. Half were asked about a job needing an introvert.

Both groups recalled real events. Both groups recalled accurately, in the sense that they were not making things up. They simply recalled different real events, the ones that fitted the question they had been given, and both groups concluded Jane was well suited to the job in front of them.

Nothing was invented. The sampling did all the work.

Fyock and Stangor found the same process maintaining stereotypes [20]. Behaviour consistent with an expectation was recalled better than behaviour that contradicted it, so the expectation kept receiving support from a memory that had already been filtered by the expectation.

If this sounds like it should be impossible, it is only impossible under a model of memory as storage and playback. Memory is not storage and playback. It is reconstruction, assembled at the moment of retrieval from fragments plus expectation, which is exactly why constructive memory produces confident recollections that never happened and why false memories form so readily in normal healthy people.

Your current belief is part of the material your brain uses to rebuild the past. So the past comes back agreeing with you. It would be strange if it did not.

Translucent glass fragments forming a glowing, curved structure in dark space.

Are the Three Really One Thing?

A reasonable question at this point: are these three things actually one thing?

It matters practically. If a single underlying tendency drives biased search, biased weighting and biased recall, then one intervention might reach all three. If they are separate, no single intervention will.

Berthet and colleagues went looking for that common factor in 2024 [21]. They adapted three classic paradigms, the Wason selection task, the 2, 4, 6 task and an interviewee task, and had 200 participants complete nine behavioural measures in total, so that each component of confirmation bias was measured in each paradigm.

The paper is unusually honest about why this is hard. It states the open question directly: whether these effects reflect confirmation bias as such, or something simpler, a general preference for positive information over negative information. Those two accounts predict very similar behaviour in most experiments and pulling them apart takes work of exactly this kind.

That question is still open. Anyone telling you confirmation bias is a single well characterised mechanism is ahead of the evidence.

The Machinery: What a Choice Does to the Evidence That Follows

Behaviour tells you what happens. It does not tell you where.

The best mechanistic study in the area is short and clean. Talluri and colleagues showed people two successive patches of moving dots and asked them to judge the direction of motion [22]. Crucially, participants committed to a judgement after the first patch, before seeing the second.

That commitment changed how the second patch was processed. Evidence agreeing with the choice already made was weighted up. Evidence disagreeing with it was weighted down. The authors describe this as a change in gain, and compare it to selective attention.

Nobody in that experiment held a belief about dots. There was no identity to protect and nothing at stake. The bias appeared anyway, within a second, purely because a choice had been made. That is a strong argument that at least part of confirmation bias sits below the level of anything we would call motivated reasoning.

The learning literature shows the same asymmetry. Palminteri and colleagues modelled how people learn from outcomes and found positive prediction errors weighted more heavily than negative ones [23]. When your choice works out better than expected you update briskly. When it works out worse you update grudgingly. They tested this for both obtained outcomes and forgone ones, and the asymmetry showed up in both.

Kappes and colleagues found the social version [24]. When another person agrees with you, you use how strongly they agree to adjust your confidence. When they disagree, you largely ignore the strength of their disagreement. Someone mildly against you and someone violently against you move you about the same distance, which is to say hardly at all. The effect tracked reduced neural sensitivity to opinion strength in the posterior medial prefrontal cortex.

There is something quietly bleak in that result. Disagreement carries information in its intensity. Someone who thinks you are badly wrong is telling you more than someone who thinks you are slightly wrong, and that extra signal is the part that goes missing.

Underneath all three results is a simpler asymmetry, and Sharot and Garrett put it at the centre of how beliefs form: whether information counts as good news or bad news changes how much of it gets through, independently of how true it is [25]. Westen and colleagues had seen an extreme version years earlier, when people evaluating politically threatening claims produced a pattern of brain activity that looked nothing like the pattern for working through a neutral problem [26].

ConfidenceDecisionEvidenceConfidenceDecisionEvidenceFirst sample arrivesCommit to a judgementRaise gain on agreeing inputLower gain on conflicting inputApparent support accumulatesResistance to revision rises

None of this is a person deciding to be unfair. It is closer to a thermostat with an asymmetric response curve.

Confidence Is the Switch

If you want one finding to carry away from this article, it might be this one.

Rollwage and colleagues used magnetoencephalography to watch what happens after a decision, and specifically what happens when the decision was made with high confidence [27]. High confidence did not simply make people less willing to change their minds. It changed what their brains did with the incoming evidence.

When confidence was high, the integration of confirming evidence was amplified. The processing of disconfirming evidence was abolished.

Not reduced. Abolished, in their word. That is one study and the word is theirs rather than mine, so hold it loosely.

That is a physical account of a thing you have probably watched happen in an argument. The moment someone becomes certain, contrary information stops landing. It is not that they are hearing it and rejecting it. Downstream of high confidence, there is much less to reject, because the signal is not being carried forward.

Which suggests something uncomfortable about certainty as a mental state. Confidence feels like the reward for having thought carefully. Mechanically it behaves more like a gate closing. It is the same architecture behind the illusion of knowing, where the feeling of understanding runs ahead of the understanding itself.

Keep that in mind for the section on fixes. It is going to matter.

Is It a Flaw, or Arithmetic You Cannot Afford?

Back to the question Klayman and Ha opened. If the behaviour is not simply irrational, what is it?

Pilgrim and colleagues published a model in 2024 that gives the most satisfying answer available [28]. Their starting point is that when you receive information, you are never updating just one thing. You are updating your belief about the claim and your belief about how reliable the source is, at the same time, and those two updates depend on each other.

Do that properly, as a full Bayesian would, and the dependencies multiply. Every belief you hold becomes entangled with every source you have ever weighed. Tracking all of it is not a matter of trying harder. It is a memory demand no real system can meet.

Their proposal is that human cognition solves this by assuming independence between beliefs when it is not actually there. It is an approximation. It makes the arithmetic possible. And confirmation bias falls out of the approximation as a side effect, without anyone needing to want anything.

That reframing is worth sitting with. On this account the bias is not a defect bolted onto reasoning. It is the visible edge of a shortcut that makes reasoning affordable.

Rollwage and Fleming argued something compatible from another direction [29]. Preferential weighting of confirming evidence can be adaptive when your confidence is well calibrated, because a decision you have good reason to trust should be somewhat resistant to noisy new evidence. The problem is not the weighting. The problem is what happens when the confidence is wrong, because then the same mechanism protects an error with the same efficiency it would protect a correct judgement.

Mercier and Sperber came at it from evolution [30]. Their argument is that reasoning did not evolve to help a solitary thinker reach truth. It evolved for argument, for producing reasons and evaluating other people's. Under that account a bias toward your own side is not a malfunction of reasoning, it is reasoning doing its job, and the correction is supposed to come from other people rather than from inside your own head.

Three quite different arguments, and all three land in a similar place. The thing we call a bug may be a feature running outside its design conditions.

Glowing indigo web with floating nodes and simplified lines.

Hot or Cold: Does Wanting Something Change What You See?

There is a long running disagreement in this field that you should know about, because a lot of confident public writing pretends it is settled.

The cold account says confirmation bias is a processing artefact. Positive test strategies, memory reconstruction, gain changes after a choice. No motives required, and the dot motion study fits this account perfectly.

The hot account says motivation does real work. Kunda made the case in 1990 [31]. People do not believe whatever they want, because they need to be able to justify the conclusion to themselves. But within that constraint, wanting a conclusion changes which evidence gets searched, which memories get retrieved and which rules of inference get applied.

The most quoted evidence for the hot account is Kahan and colleagues on motivated numeracy [32]. Participants were given a two by two table of numbers and asked to draw the correct conclusion. When the table was about the effectiveness of a skin cream, the more numerate people did better, as you would expect. When the identical numbers were relabelled as being about gun control, the more numerate people did worse when the correct answer conflicted with their politics. Higher skill produced a bigger gap, not a smaller one.

The shape of that result is what makes it hard to dismiss. If skill simply failed under pressure you could call it a performance problem. Skill did not fail. It was recruited, and it went to work on the wrong side.

Van Bavel and Pereira built a wider account around results like that one [33]. Their argument is that a political belief is not only a claim about the world, it is also a membership badge, and the cost of getting the claim wrong is usually far smaller than the cost of losing the badge.

Now the other side, which is substantial.

Pennycook and Rand argue that a lot of what looks like motivated reasoning is better explained as insufficient reasoning [34]. Their data on belief in fabricated headlines pointed the same way for people across the political spectrum: those who engaged more analytic thinking were better at telling true from false, regardless of whether the headline flattered their side. Lazy, in their framing, rather than biased.

Tappin, Pennycook and Rand raised a sharper methodological objection [35]. Most partisan bias studies are correlational in a way that makes the causal claim unsafe. People who hold a prior belief also hold different background knowledge and different trust in sources, so evidence that they update differently does not by itself show that motivation caused the difference. They followed up with experimental work arguing that the link between cognitive sophistication and politically motivated reasoning is weaker and more fragile than the headline claims suggest [36].

This objection is not a technicality. A great deal of public writing about political bias treats these studies as showing that people cannot process facts that threaten them. If the design cannot separate motivation from background knowledge, that conclusion is doing more work than the data supports.

And Ditto and colleagues ran a meta analysis on a question that generates more heat than any other in this area, which is whether one political side is more biased than the other [37]. Their answer, in their own phrasing, is that at least bias is bipartisan. They found the effect comparably sized on both sides.

The honest summary is that hot and cold accounts both have evidence, they are not mutually exclusive, and anyone who tells you which one is correct is telling you their opinion.

Myside Bias, and Why Being Clever Does Not Help

There is a closely related term you will meet and it is worth keeping separate.

Myside bias is the tendency to evaluate evidence, generate evidence and test hypotheses in a way that favours your own prior opinions. Stanovich, West and Toplak have argued that it behaves differently from other reasoning biases in one specific and troubling way [38]. Most reasoning failures correlate with cognitive ability. Better reasoners make fewer of them. Myside bias largely does not follow that pattern.

Being clever does not protect you. It gives you better tools for defending what you already think.

Wolfe and Britt located where this shows up in written argument [39]. When people write a persuasive piece, whether they include and address the other side depends heavily on what they believe a good argument is supposed to look like. Many people hold a model in which acknowledging opposition is a weakness rather than a strength, and they write accordingly.

Stanovich developed the full argument in a 2021 book [40], where he treats myside bias as a distinctive problem precisely because the usual remedies, more education and more intelligence, do not touch it.

TermWhat it meansWhen it firesWhere the name comes from
Confirmation biasSeeking weighing and recalling in line with a current beliefAcross search reading and recallMynatt and colleagues 1977
Myside biasFavouring your own prior opinion when generating and testing evidenceWhen you argue or evaluate a claim you holdStanovich and colleagues
Motivated reasoningWanting a conclusion shapes which evidence and rules get usedWhen the answer has something at stake for youKunda 1990
Selective exposureChoosing which material to look at in the first placeAt the moment of choosing what to readFrey 1986
Belief perseveranceThe belief survives after its evidence is withdrawnAfter a retraction or a debriefingRoss Lepper and Hubbard 1975
AnchoringAn early value pulls every later estimate toward itAt the first number or first diagnosisDistinct from confirmation bias and easy to confuse with it

The table above matters more than it looks. A great deal of confused writing about bias comes from treating these six terms as synonyms, and they are not. They fire at different moments and they need different countermeasures.

What It Costs in a Consulting Room

Abstractions are easy to shrug off. Consequences are not.

Mendel and colleagues built a clean test of confirmation bias in diagnosis [41]. They gave a case to 75 psychiatrists and 75 medical students. Each participant formed a preliminary diagnosis, then chose which further pieces of information to look at from a set that included both confirming and disconfirming material.

Thirteen percent of the psychiatrists and twenty five percent of the students searched confirmatorily. And the part that matters: those who did were significantly less likely to arrive at the correct diagnosis.

Two things deserve attention there. The first is that most participants did not show the bias, which is a useful corrective to the impression that this happens to everyone all the time. The second is that the psychiatrists were roughly half as likely as the students to fall into it. Tempting as it is to call that the effect of experience, this design cannot show it. Psychiatrists and students differ in training and in age and in how much they already know, and a comparison between two groups does not separate those out.

Croskerry argued in 2003 that diagnostic error should be treated as a reasoning problem and not only a knowledge problem, and that medicine had spent decades teaching content while ignoring the process the content gets used in [42]. Graber and colleagues then went through real diagnostic errors in internal medicine to see where they actually came from [43]. Their conclusion was that most errors involved several contributing factors at once rather than one clean cause, with cognitive factors and system factors tangled together. This is the mechanism sitting behind why fast thinking produces wrong diagnoses.

But the field disagrees with itself here, and the disagreement is worth reporting.

Norman and colleagues argued in 2017 that the cognitive bias account of diagnostic error has been overstated [44]. Their case is that errors correlate more strongly with gaps in knowledge than with failures of reasoning style, and that the debiasing programmes built on the bias account have not delivered what was promised. If a doctor does not know a disease, no amount of slowing down will conjure it.

Martínez and colleagues added a measurement problem in 2026 [45]. In the vignette studies that most of this evidence rests on, anchoring and confirmation bias are extremely difficult to tell apart, because the design that produces one also produces the other. That does not mean the effects are not there. It means some of the published effect sizes are measuring a blend.

This is a good example of a field getting stricter with itself rather than louder. The clinical effects are almost certainly real and they matter. Some of the published numbers are still measuring two things at once, and saying so is how the estimates eventually get better.

Share showing confirmatory search in a diagnostic taskPsychiatristsStudents302826242220181614121086420Percent of participants

Five Experts and One Fingerprint

The single most unsettling study in this whole literature takes about two minutes to describe.

Dror, Charlton and Péron took a pair of fingerprints that five latent print experts had examined years earlier in a real criminal case and declared a match [46]. Between them those five had eighty five years of experience. The researchers put the same prints in front of the same experts again, this time with context suggesting these were the prints behind a notorious wrongful arrest.

Four of the five reversed themselves. Three now said the prints did not match. One said no definite decision could be made. Only one expert still called it a match.

Five people is a very small study and it should be read as one. But notice what it is a small study of. Not students guessing. Experts re-examining their own past conclusion, on physical evidence that had not changed by a single ridge, and reaching a different answer because the story around it was different.

Fingerprint comparison feels like reading a measurement off an instrument. It is not. It is a perceptual judgement made by a person who knows things about the case, and what they know reaches into the judgement.

Kassin, Dror and Kukucka gave this a name and a framework in 2013 [47]: forensic confirmation bias. Their central observation is that the contamination usually runs in one particular direction. Forensic examiners often know things about the case before they look, and those things are rarely neutral.

Kukucka and Kassin demonstrated the same effect with handwriting evidence [48]. People told that a suspect had confessed judged handwriting samples as more similar than people who were told nothing.

Dror later set out six fallacies that make expert decision makers believe they are immune, most of which amount to variations on the idea that expertise protects you [49]. The Mendel result suggests expertise does help. It clearly does not immunise.

The response in forensic science has mostly not been to train examiners harder. It has been to change the workflow so the context never reaches them. That distinction is going to be the point of the last part of this article.

The same thing happens once a legal case has a theory attached to it. An early account of what happened starts filtering what counts as evidence for everything that follows, and belief perseverance keeps that account standing after its support has gone [50].

Intricate pale gold contour lines on deep indigo, glowing red misalignment.

Why You Cannot Feel It Happening

There is a reason all of this is so hard to act on, and it is not that people are careless.

Confirmation bias produces no sensation. There is no moment where you notice yourself skipping the study that disagrees. The biased search feels like a search. The stricter standard applied to unwelcome evidence feels like rigour, and it feels most like rigour when it is working hardest, because spotting the flaw in a weak paper is genuinely a skill. The filtered recall feels like remembering.

Every stage is invisible from the inside precisely because every stage is doing something that would be correct in a slightly different situation.

Which is why the bias is so much easier to see in other people. When Heerma van Voss and colleagues studied professional risk analysts they measured two things at once: confirmation bias itself, and the blind spot that travels with it, meaning the gap between how biased you judge yourself to be and how biased you judge everybody else to be [51]. That gap is why a room can agree unanimously that confirmation bias is a serious problem and every person in it can mean somebody else.

There is a particular trap in reading an article like this one. Every example above arrives with the answer attached. You know the rule was any ascending numbers, you know the two studies were fabricated and matched, you know the fingerprints were the same prints. In each case the correct behaviour is obvious because you are standing outside the situation with the design in front of you.

The participants were inside it. They had a hypothesis that seemed reasonable, evidence that seemed to support it, and no marker anywhere telling them a test was running.

That asymmetry is not a minor detail. It is most of the phenomenon. Any strategy that depends on you noticing in the moment is a strategy that has assumed away the thing it is supposed to solve, and it is the single best argument for the approach in the last section of this article: change the situation, because the person cannot be relied on to catch it while inside one.

The Claim That Would Not Replicate

Now for a piece of self correction, because an article about believing things too easily should demonstrate what changing your mind looks like.

You have probably heard the backfire effect. It says that correcting a false belief can strengthen it, that showing someone evidence against their position pushes them further in. It appears constantly in journalism, in training materials, and in advice about how to talk to people who disagree with you.

Nyhan and Reifler reported it in 2010 [52]. It is a striking result and it spread fast, partly because it fits a satisfying story about how hopeless argument is.

Then people tried to reproduce it.

Wood and Porter ran a large scale attempt across many political issues and thousands of participants [53]. Their title tells you the result: the backfire effect is elusive. Corrections generally moved people toward accuracy, including on issues where their side was implicated. They found the effect very hard to produce.

Guess and Coppock tested whether counter attitudinal information causes backlash across several experiments and found no support for it either [54].

The most notable response came from Nyhan himself. In 2021 he published a paper arguing that the backfire effect does not explain the durability of political misperceptions [55]. Misperceptions do persist, but backfire is not why, and the reasons involve directional motivations, source trust and how rarely people encounter corrections at all.

That is what a healthy field looks like. A striking result, an attempt to reproduce it, a failure, and one of the original authors publishing the correction.

The practical lesson is genuinely good news. Corrections mostly work a little. They do not usually make things worse.

Lewandowsky and colleagues had already mapped why corrections nonetheless underperform [56]. Retracted information keeps influencing reasoning even when people accept the retraction, because the original claim is woven into a causal story and removing it leaves a hole the mind fills. Ecker and colleagues brought that account up to date in 2022, and what is striking is how little of it involves backfire [57]. It is about how rarely a correction is seen at all, how much the source is trusted, and how firmly the false claim has been built into somebody's explanation of events. Pennycook and Rand arrive somewhere similar from the fake news side [58].

Siebert and Siebert tested what actually reduces belief perseverance after a retraction, and two things helped: warning people in advance that retracted claims tend to linger, and pairing the retraction with a direct counter argument instead of a bare withdrawal [59]. Notice the shape of that. Neither one is the correction working harder. Both change what surrounds it.

So: correct people. It usually helps a bit. Just do not expect the correction to remove the belief the way deleting a file removes a file.

The Algorithm Is Not the Whole Story

Here is where I had to change my own mind while researching this, so I will show the working.

The popular account is that recommendation algorithms sort us into echo chambers, feeding each of us a diet of agreement until we cannot see the other side. Confirmation bias, industrialised. It is a tidy story and I expected the evidence to support it.

It mostly does not.

Bakshy, Messing and Adamic studied 10.1 million United States Facebook users and separated three stages: who your friends are, what the ranking algorithm shows you, and what you click [60]. All three reduce exposure to cross cutting content. But individual choice reduced it more than the algorithm did. The strongest filter in the chain was the person. That study drew criticism when it appeared, largely because it could only look at the minority of users who had declared a political affiliation, so treat it as the weakest of the four results in this section rather than the anchor.

Cinelli and colleagues compared echo chamber structure across four platforms using more than 100 million pieces of content [61]. The important word there is compared. Echo chamber effects were much stronger on some platforms than others, which means the phenomenon depends on how a platform is built rather than being an inevitable consequence of algorithms in general.

Then came the experiments, and they are the strongest evidence available.

Nyhan and colleagues, working with data covering the entire population of active adult Facebook users in the United States, ran a field experiment on 23,377 people during the 2020 election [62]. For those users they cut exposure to content from like minded sources by about a third. Like minded content is indeed the majority of what people see. Reducing it did not produce the attitude changes the echo chamber account predicts.

Guess and colleagues, in the same research programme, moved consenting users off the algorithmic feed entirely and onto a reverse chronological one [63]. What people saw changed substantially, including seeing more political content and more untrustworthy content. What did not change: issue polarization, affective polarization and political knowledge.

And Törnberg and colleagues showed by modelling that echo chambers can emerge with no algorithmic personalization at all and no preference for similar others [64]. Structure alone can produce them.

None of this means platforms are harmless or that design does not matter. It means the specific claim, that algorithmic curation is what makes people close minded, is not well supported, and that the older and less convenient explanation, that we do this to ourselves and always did, keeps surviving contact with data. That pattern repeats whenever a technology gets blamed for a habit older than it is. The evidence on whether AI tools damage critical thinking has the same shape: the loudest fears are the least supported, and the real effects are quieter and much harder to sell.

Barberá and colleagues had found earlier that the picture varies by topic [65]. Political subjects looked more segregated than non political ones, which again points at people rather than machinery.

Vast glowing network of nodes forming amber clusters in dark space.

Your Question Arrives Already Bent

There is a more interesting version of the technology story and it goes back to the first of our three stages.

Leung and Urminsky published the clearest work on this in 2025 [66]. They call it the narrow search effect and they tested it across 21 studies, 14 of them preregistered, on real platforms including Google and ChatGPT and an artificial intelligence powered version of Bing, as well as custom built search interfaces they controlled.

The mechanism has two parts and neither is the algorithm being sinister.

First, your prior belief shapes the words you type. Someone who suspects coffee is harmful searches differently from someone who suspects it is beneficial, and neither of them is trying to cheat. Second, search engines are built to be responsive. They return what matches the query. Put those together and you get a narrow set of results that reflects the question rather than the topic, and belief updating is limited as a result.

The bias is in the query, not the ranking. It arrives before the system does anything.

Which explains why their interventions are interesting. They tested both user side changes, getting people to search more broadly, and algorithm side changes, having the system return a wider range regardless of how the question was phrased. Broadening the search promoted belief updating.

The generative version of this problem is arriving now. Lopez-Lopez and colleagues describe how conversational systems mediate confirmation bias in health information seeking specifically [67]. Their concern is hypercustomisation: a system that adapts closely to how you phrase things, and to what you seem to want, can reflect your assumptions back at you with more fluency and more apparent authority than a list of links ever could. They identify pressure points where the bias enters, and the first is where you would expect by now. Query phrasing.

You cannot outsource the framing of the question. Whatever answers it will answer the question you asked.

What Actually Works, and What Only Sounds Like It Does

This is the section most articles get wrong, because it is where the reassuring listicle lives.

Start with the technique that everybody recommends. Lord, Lepper and Preston tested it in 1984 [68]. Consider the opposite: before judging, actively ask yourself what you would think if the evidence pointed the other way. It worked in their study, and a simple instruction to be unbiased did not.

That is a real result. But it is one study from 1984, and it has been carrying an enormous amount of popular advice for forty years.

Whitt and colleagues tested three debiasing approaches directly against confirmation bias in 2023 [69]. A social norms technique, telling people that open minded information seeking is what people like them do, reduced selective exposure relative to control. Consider the opposite showed little evidence of working. Psychoeducation, explaining the bias to people, showed little evidence of working either.

Read that again, because it is the finding that should change your behaviour. Explaining confirmation bias to someone is among the weakest interventions tested. Which has an awkward implication for articles like this one.

Before concluding that nothing works, look at what separates the failures from the successes. Psychoeducation is explanation. You get told the bias exists and you are left to do something with that. The things that do move the needle are not explanations at all. They are procedures, a specific action to perform at a specific moment, and that distinction runs through everything below.

Heerma van Voss and colleagues studied confirmation bias in national risk analysts, the professionals who assess threats like pandemics and conflict for a European government, alongside a matched sample of masters students [51]. Two findings. The analysts showed less confirmation bias than the students, both on risk related judgements and on unrelated ones, so professional training and selection appear to do something. And a one shot debiasing training session reduced confirmation bias in both groups.

Branchini and colleagues took the fix right back to Wason's rule discovery task [70]. They compared three conditions: no prompt, a prompt to analyse the properties of the starting triple, and a prompt to analyse those properties and then identify their opposites. Thinking in opposites nearly doubled the success rate and produced more first attempt discoveries of the rule.

The mechanism they report is the useful part. Success did not come from testing more triples. It came from people repeating the same hypothesis less, and noticing the dimension that actually mattered.

Transfer is the question that matters, and there is at least one good answer to it. Sellier, Scopelliti and Morewedge trained people with a game, then measured them weeks later on a real business decision in a setting that looked nothing like the game [71]. The training still showed up. That built on earlier work from Morewedge and colleagues in which a single session produced improvement lasting across several different biases [72]. Lilienfeld and colleagues had already argued this work should be treated as something to give away rather than something to publish [73], which is a fair description of where it still is not.

So where does that leave you?

Awareness alone does very little. Both sides of the effectiveness debate agree on that much. It is the least controversial claim in this article.

Specific trained procedures do more than general exhortation. Thinking in opposites beats trying to be objective, and it beats it because it gives you something concrete to do.

And the interventions with the best record change the situation rather than the person. Blind protocols in forensic labs. Pre-specified inclusion criteria before a review begins. Pre-registration before data collection. Structured search that broadens the query for you. These work because they do not depend on anyone noticing they are biased in the moment, which is precisely when nobody notices.

Person

Person

Situation

Situation

Want less biased judgement

Person or situation?

Explain the bias

Trained procedure

Blind the context

Broaden the search

Little measured effect

Works when specific

There is one more thing worth knowing, and it connects back to the confidence finding.

Rollwage and Fleming argued that whether confirmation bias hurts you depends on how good your metacognition is [29]. If your confidence tracks your accuracy, then weighting confirming evidence more heavily is reasonable behaviour. If your confidence is poorly calibrated, the same mechanism defends your errors as efficiently as it defends your correct beliefs.

That makes calibration the place to push rather than bias itself. Knowing how much to trust your own judgement is a trainable skill, and it is the practical core of what metacognition does for learning. Get better at knowing when you are probably wrong and the same machinery starts working for you.

And remember the group finding, because it cuts against the obvious workaround. Groups searched more selectively than individuals. Assembling a team does not fix this unless the team is built so that dissent is cheap for the dissenter, which most teams are not.

What This Leaves You With

We started with three numbers and a rule you did not find.

The standard telling of that story is that people are hopeless at testing their own ideas. The more accurate telling is that people use a strategy which usually works, in a task built so that it could not, and that being told the strategy is a bias does not by itself make anyone better at anything.

Everything else follows from taking that seriously.

It is not one thing. It is three, and they fire at different moments in different parts of the machine. It turns up in judgements about moving dots, where nobody has anything at stake at all, so it cannot all be about motives. It may be what careful updating looks like once you admit that the careful version costs more memory than anyone has. Confidence does not simply make people stubborn. On the evidence available it seems to gate contrary information out before there is anything left to be stubborn about. Expertise helps and does not save you. And the technology story people tell most often, the one where the algorithm builds your bubble, keeps failing its own experiments, while a quieter version of it holds up perfectly well: the bias was already in the question you typed.

The uncomfortable part is what follows for advice. Explaining the bias barely helps. Deciding to be objective barely helps. What helps is specific: a concrete procedure like generating the opposite, a workflow that keeps the contaminating context away from the judgement, a search that gets broadened whether or not you thought to broaden it, and better calibration of your own confidence.

None of that is inspiring and all of it is boring to implement. That is usually a sign something is real.

So here is the one thing worth doing differently, and it needs care, because the obvious version of it is the version the evidence just knocked down. Resolving to consider the opposite, as an attitude you carry into reading, has a poor record.

What has a better record is changing the input before any of the machinery starts.

Before you look something up, write one line saying what result would change your mind. Then search for that line, in those words.

The point is not the resolve. The point is that the sentence becomes the query, and a query aimed at what would refute you returns a different page of results than a query aimed at what you already suspect. You are not trying to stay fair minded while you read. You are changing what arrives.

That is the 2, 4, 6 problem in miniature. Wason's players did not need better character. They needed to type one triple they expected to fail.

Frequently Asked Questions

What is confirmation bias in simple terms?

It is the tendency to look for weigh and remember information in a way that fits what you already think. It is not lying to yourself and it is not stupidity. It happens in three separate places: in how you frame a question before you see any evidence, in the standard you apply to evidence once it arrives, and in which episodes come back when you recall something days later. Because the three stages are separate, a technique that helps with one of them often does nothing for the other two. That is a large part of why the usual advice underperforms.

Who discovered confirmation bias and what was the original experiment?

Wason ran the founding experiment in 1960. He gave people the number triple 2 4 6 told them it followed a rule and let them test any other triples they wanted. Most people guessed ascending by twos tested only triples that fitted that guess collected a run of yeses and announced the wrong rule. The real rule was any three ascending numbers. The phrase confirmation bias itself came later from Mynatt and colleagues in 1977. What is usually left out is that Klayman and Ha showed in 1987 that testing cases you expect to be positive is normally an efficient strategy and only fails when your hypothesis sits entirely inside the true rule which is exactly how Wason built the task.

Does telling people about confirmation bias reduce it?

Barely. This is the most awkward finding in the area and it applies to articles like this one. Whitt and colleagues tested three approaches against confirmation bias in 2023 and found little evidence that psychoeducation or consider the opposite worked while a social norms technique did reduce selective exposure. What does better is specific and procedural rather than general. Branchini and colleagues found that prompting people to identify the opposites of the properties they had noticed nearly doubled the success rate on Wason's task. Heerma van Voss and colleagues found a single training session reduced confirmation bias in professional risk analysts and in students. Interventions that change the situation rather than the person have the best record of all.

Is the backfire effect real?

The evidence says mostly no. The backfire effect is the claim that correcting a false belief pushes people further into it. Nyhan and Reifler reported it in 2010 and it spread widely. Wood and Porter went looking for it across many issues and thousands of participants and could not reliably produce it and their paper is titled the elusive backfire effect. Guess and Coppock found no support for backlash either. Nyhan himself published a paper in 2021 arguing that the backfire effect does not explain why political misperceptions persist. Corrections usually move people a little toward accuracy. They rarely make things worse. Misperceptions do persist but for other reasons.

What is the difference between confirmation bias and myside bias?

Confirmation bias is the broad tendency to favour information that fits a current hypothesis and it appears even when nothing is at stake for you. Myside bias is narrower and is specifically about favouring your own prior opinion when you evaluate evidence or build an argument. The distinction matters because of one finding. Most reasoning failures get smaller as cognitive ability rises. Myside bias largely does not. Stanovich and colleagues have argued that being clever gives you better tools for defending what you already believe rather than better protection against believing it too easily.