The Number That Should Not Have Mattered

In a psychology laboratory, a group of experienced legal professionals read a criminal case file. Then they rolled dice.

The roll produced a number. Each professional was then asked whether the sentence the prosecutor had demanded, a figure tied to that roll, was too high or too low. And then they were asked what sentence they themselves would impose.

The people who rolled high recommended longer sentences than the people who rolled low [1].

These were not undergraduates. They had years of courtroom experience. They knew, because everyone knows, that a number produced by dice has nothing to do with what a defendant deserves. It moved them anyway.

That study is one of the most quoted demonstrations of what psychologists call the anchoring effect: the tendency for a number you encounter first to pull every judgment that comes after it toward itself. You have almost certainly read about it. It turns up in negotiation courses, in pricing seminars, in the standard list of cognitive biases, and in a few thousand articles that all cite the same 1974 paper and then tell you to do your research before you shop.

This article is going to tell you something those articles do not, and it is not what you expect.

The anchoring effect is in good health. Fifty years of testing, an era of replication failures that flattened a long list of famous psychology findings, and it is still there. In 2026 researchers stood people on a platform in a virtual world that could be set at any height, and asked them how high up they were. Even with the whole scene visible around them to judge from, an arbitrary number they had seen first moved their answers [2].

What is not in good health is the explanation. The mechanism you will find in almost every textbook and almost every page ranking for this topic has been failing its own laboratory test since 2019. And the particular version of anchoring that consumer articles like best, the one where a number you merely noticed in the room changes what you would pay for dinner, has failed to replicate more than once.

So the honest version of this story has two halves that most writing on the subject refuses to hold at the same time. The effect is real. The account of why it happens is unsettled, and one popular branch of it is probably wrong.

Keeping those two halves together is harder than it sounds, and it is where most coverage gives up. Drop the first half and you get a debunk. Drop the second and you get the same article everyone else has written. The rest of this piece is an attempt to hold both: first the mechanism argument in the order it happened, then how far the surviving version of the effect actually reaches.

Heavy iron anchor in deep blue-green water with drifting sediment.

The Wheel and the United Nations

The founding experiment is almost absurdly simple, which is part of why it became famous.

Amos Tversky and Daniel Kahneman sat participants in front of a wheel of fortune marked from 0 to 100 and spun it. The wheel stopped on a number. Participants were then asked two questions in sequence. First: is the percentage of African countries in the United Nations higher or lower than that number? Second: what is the actual percentage?

People who saw a high number gave higher estimates than people who saw a low number [3].

Nobody in that room had any reason to believe a spinning wheel knew something about the composition of the United Nations. The number arrived in front of them, attached to nothing, carrying no information about the question. It moved their answers regardless.

Tversky and Kahneman published this in Science in 1974, inside a paper that also introduced availability and representativeness. They were not arguing that people are stupid. Their argument was that judgment under uncertainty runs on shortcuts, that the shortcuts usually work, and that you can see the machinery most clearly in the cases where they fail. Anchoring was one such case. You start from whatever value is in front of you and you adjust from it, and you stop adjusting too soon.

It is worth saying that the word did not arrive with them. Sixteen years earlier, Muzafer Sherif and colleagues had used "anchoring stimuli" to describe how a reference point shifts perceptual judgments, producing assimilation toward the anchor in some conditions and contrast away from it in others [4]. The 1974 paper moved the idea out of psychophysics and into decisions about money, risk and other people. That is the leap that mattered.

It is a leap worth noticing, because it changed what the finding was for. A shift in how you judge the weight of a block is a curiosity. A shift in how you judge what a house is worth, or how long a sentence should be, is something else.

Twenty-one years later, Karen Jacowitz and Kahneman formalised the procedure into something other laboratories could copy exactly [5]. Ask a comparative question against an anchor. Then ask for an absolute estimate. Compare the estimates across high-anchor and low-anchor groups, and express the gap as an index.

That two-step recipe is now called the standard anchoring paradigm, and it is the reason this literature is comparable across decades. It is also, as you will see, the version that survives.

A review counted the field at over a hundred published studies by the time it was written, spread across pricing, forecasting, medicine, law and negotiation [6]. By any normal standard the phenomenon was settled. What was not settled, and still is not, is why it happens.

What an Anchor Actually Does to You

Before the argument about mechanism, it helps to be precise about what is being claimed.

An anchor is a number that enters your head before you produce a number of your own. That is all. In the standard paradigm it does not have to be relevant, it does not have to be plausible, and it does not have to be believed. How far that tolerance extends is one of the things this article is about, and the answer turns out to be less far than the popular version claims.

What it does is give you a starting point. And a starting point is not a neutral thing, because judgment under uncertainty is not a lookup. You do not know what a house is worth. You do not know how many countries are in anything. What you have instead is a fuzzy region of plausible answers, and something has to pick a spot inside it.

The anchor picks the spot. Then you move.

The whole argument of the last thirty years is about what happens during that move. There are two families of answer, they were developed by different research groups, and for a long time the field treated them as rivals. Both are still standing, and neither has won cleanly.

It is worth pausing on why this is hard to study, because the difficulty explains a lot of the mess that follows. You cannot ask people. Anchoring does not appear to be available to introspection. There is a study later in this article where professionals were asked directly what they had used to reach a number, and the anchor was not on their list. So researchers have to infer the mechanism from indirect signatures: how fast a related word is recognised, how far the estimate lands from the anchor, what happens when you add time pressure or take it away. Each signature is a proxy. When a proxy stops behaving, you cannot immediately tell whether the theory is wrong or the proxy was never measuring what you thought.

That is exactly the situation the field found itself in, and it is the reason the honest answer to "why does anchoring happen" is longer than one sentence.

The First Explanation, and the Test It Keeps Failing

In 1997 Fritz Strack and Thomas Mussweiler published an account called selective accessibility [7]. It is elegant, and once you have understood it you will see it everywhere in the popular coverage, usually without attribution.

The idea starts from the comparative question. When you are asked whether the Elbe is longer or shorter than 890 kilometres, you do not shrug. You try to answer it. And the cheapest way to answer it is to test the proposition that the river is about 890 kilometres long, which means going looking for everything you know that fits with a long river.

By the time you have finished, your head is full of long-river facts. Then someone asks you how long the Elbe actually is, and you answer from what is currently accessible. Which is long-river facts.

On this account the anchor is not a starting point you adjust away from. It is a search term. It changes what you retrieve, and you then answer honestly from a biased sample of your own knowledge.

Mussweiler and Strack built a specific laboratory test for this. If the comparative question really does activate anchor-consistent concepts, those concepts should be easier to process immediately afterwards. So they gave participants a lexical decision task, where you see a string of letters and press a key to say whether it is a real word. Words related to the high-anchor scenario should be recognised faster by people who saw the high anchor [8].

That is a semantic priming signature, and it was found.

That result became the signature test for selective accessibility. If you wanted to argue that a particular anchoring effect ran on this mechanism, you showed the priming. The account was extended across the following years to cover how comparison drives judgment more generally [9].

Then in 2019, a group led by Adam Harris tried to use it, and could not make it work.

They ran four experiments with the lexical decision signature test, on temperature estimates and on car prices, powered properly. Each experiment was designed so that every individual comparison had at least an eighty percent chance of detecting a real effect, which required sixty-three participants in each condition and a hundred and twenty-six per experiment. Experiments one through three ran a hundred and twenty-eight native English speakers each. None of those four found the priming signature. A fifth experiment swapped in a continuous identification task, a more sensitive measure of processing fluency than lexical decision, with a hundred and twenty-eight participants. It also found nothing [10].

Here is the part that matters, and that almost nobody reports.

The anchoring effects themselves were there. In the same experiments, on the same participants, the anchors moved the estimates exactly as they were supposed to. What was missing was the fingerprint of the mechanism. The authors' own conclusion is that the sturdiness of anchoring effects is remarkable, and that the theoretical basis for these particular tests is shaky, and they advise against using the test for this purpose.

That is a careful sentence and it deserves a careful reading. It is not "selective accessibility is refuted". It is "the standard way of demonstrating selective accessibility does not demonstrate anything". Those are different claims, and the difference is the difference between a field correcting itself and a field collapsing.

Trouble had been visible from inside the camp, too. Bahník and Strack found a condition where the effect goes away: when the information accessible for the comparison overlaps heavily with the information needed for the estimate, the anchoring effect is undermined rather than strengthened [11].

A theory that predicts more anchoring from more overlap has a problem when the data say the opposite.

More recently, work using item-based anchoring found something that no simple one-directional account expects: judgments of both items assimilate toward each other, rather than one item dragging the other [12].

The pull is mutual. That is not a small detail. It suggests the process is closer to a comparison finding its own middle than to a starting value being reluctantly abandoned.

The Second Explanation, and Its Own Complication

The other family of answers is older, because it is what Tversky and Kahneman originally said: you start at the anchor and you adjust, and you stop before you should.

For a while this was treated as obviously wrong. Study after study failed to find evidence of adjustment when anchors were provided by the experimenter. Then in 2001 Nicholas Epley and Thomas Gilovich pointed out that everyone had been testing the wrong kind of anchor [13].

Their split is the most useful single idea in this literature, so here it is slowly.

Some anchors are handed to you. A listing price, an opening offer, a number on a spinning wheel. You have no reason to think the true answer is near them, so you have no reason to start there and creep away.

Other anchors you generate yourself. Asked what year George Washington was elected president, an American does not draw a blank. They think: 1776, the Declaration. Then they adjust upward, because they know the election came later. Asked the freezing point of vodka, they think: water freezes at zero, alcohol is lower than that, so somewhere below.

For self-generated anchors, adjustment is real, observable and insufficient. Epley and Gilovich showed that people adjust from these starting points, that the adjustment stops early, and that anything which reduces mental effort makes it stop even earlier [14]. The distinction is not academic. It predicts who you can anchor, on what, and how hard you would have to push. Their later summary makes the two-process picture explicit: different anchors, different machinery, one label [15].

For roughly a decade this looked like the settlement. Two mechanisms, one for each kind of anchor, everybody goes home.

Then Joseph Simmons and colleagues showed that accuracy motivation does move estimates on provided anchors, in the direction of the correct answer, which is exactly what the clean split says should not happen [16].

Not a refutation. A complication, and an honest one.

Individual differences add another wrinkle. When you look at who is most affected, the pattern fits insufficient adjustment better than it fits pure accessibility [17]. And Gretchen Chapman and Eric Johnson had been mapping the limits of the whole thing from early on, finding both where anchoring stops and how anchors shape which features of a target get weighted at all [18] [19].

So where does that leave the mechanism question? Roughly here: two accounts, each with real support, each with results the other explains better, and a signature test for one of them that no longer works. A recent chapter-length synthesis by researchers inside the debate treats multiple mechanisms as the working assumption rather than a failure to choose [20].

Given to you

You produced it

A number arrives

Source of the number?

Comparative question

Known starting fact

Anchor-consistent knowledge activated

Adjust away from it

Estimate pulled toward anchor

Notice what the diagram does not say. It does not tell you which branch is correct, because the field does not know. What it tells you is that two quite different things are being called by one name, which is the source of most of the confusion in the popular coverage.

Three Different Things Wearing One Name

If you take nothing else from this article, take this. "Anchoring" refers to at least three separable phenomena, and they are not in the same evidential condition.

The first is standard anchoring, where you are explicitly asked to compare a quantity against a number and then estimate it. This is the Jacowitz and Kahneman procedure. It is the one in the dice study, the wheel study, and most of the laboratory literature.

The second is self-generated anchoring, where you produce your own starting value and adjust from it.

The third is basic or incidental anchoring, where there is no comparative question at all. A number is simply present, and it allegedly bends your judgment anyway. Timothy Wilson and colleagues argued in 1996 that this works, that mere exposure to a number is enough provided you have paid it some attention [21].

Eleven years later Clayton Critcher and Thomas Gilovich took it further into everyday life. Their studies included the one every pop-psychology page reaches for: people estimated they would spend more at a restaurant called Studio 97 than at one called Studio 17 [22].

A meaningless number in the name of a business, changing what you think dinner costs. It is a wonderful finding. It is the finding that has not held up.

Kind of anchoringWhat the anchor isProposed mechanismHow it is holding up
StandardA number you are explicitly asked to compare againstSelective accessibility or adjustmentReplicates reliably including in 2026
Self-generatedA fact you retrieve as a starting pointInsufficient adjustmentReplicates. The clean split has been complicated
Basic or incidentalA number merely present in the environmentDisputedFailed to replicate in multiple attempts

The right-hand column is where the popular coverage and the research literature have come apart. Almost every explainer you will read describes the third row and cites evidence from the first.

The Version That Failed

Many Labs 2 was one of the largest replication projects psychology has run. Twenty-eight classic and contemporary findings, protocols peer reviewed before any data were collected, each protocol administered to roughly half of a hundred and twenty-five samples made up of fifteen thousand three hundred and five participants across thirty-six countries [23].

Across the whole project, fifteen of the twenty-eight replications produced a statistically significant effect in the same direction as the original. Those are the project's numbers, not anchoring's, and they get quoted out of context constantly. What matters here is the specific result: the Critcher and Gilovich study two finding was among those that did not replicate.

Then in 2020, David Shanks and colleagues went after incidental environmental anchoring directly. Three studies, using the Critcher and Gilovich method, measuring consumer price estimations. They found no statistically significant evidence of incidental anchoring [24].

And in the same studies, standard anchoring came through strongly.

Read that pairing again, because it is the cleanest fact in this entire topic. Same participants, same laboratory, same measurements. The version where you are asked to compare against a number worked. The version where a number is just lying around did not.

This is not a unanimous verdict. Other groups have reported incidental effects on willingness to pay [25], and more recent work finds that spontaneously generated anchors shift how people divide quantities and how they behave afterwards [26]. The question is open. But it is open in the way a contested claim is open, not settled in the way every consumer article implies when it tells you the restaurant name is manipulating your wallet.

Meanwhile the standard paradigm keeps refusing to die. The virtual reality study mentioned at the start of this article is a good example of why. Its authors point out that almost all anchoring research is laboratory-based and gives participants no natural way to sample the information they need. So they built a virtual world with a platform that could sit at any height, showed people an anchor, and then let them judge how high up they were from what they could actually see. The anchoring persisted even though reliable information about the true answer was in front of them at the moment of judgment [2].

Sorting out which of these results generalise is now its own research problem, and the most ambitious attempt to pool fifty years of moderator findings into one picture is currently an unreviewed preprint [27]. That is a fair summary of where the field is. The question of what makes anchoring stronger or weaker is live enough that the best synthesis of it has not finished peer review.

The Decade Everyone Skipped

Put the dates in order and one stretch of this history stands out.

1974
Tversky and Kahneman publish the wheel of fortune result
1996
Wilson argues an anchor works without any comparison
1997
Strack and Mussweiler propose selective accessibility
2001
Epley and Gilovich put adjustment back for self-generated anchors
2006
Legal professionals are moved by a throw of dice
2007
Critcher and Gilovich report incidental environmental anchors
2018
Many Labs 2 fails to replicate an incidental anchoring result
2019
Five experiments fail to find the selective accessibility signature
2020
Three studies find no incidental anchoring but strong standard anchoring
2026
Anchoring persists in a virtual reality height estimation task

The first twenty years found the effect. The next twenty argued about what caused it. Everything from 2018 onward has been the field checking which parts of the pile were load bearing, and that is the stretch almost no popular coverage mentions, which is why the version most people carry around is the version the last eight years complicated.

Antique brass plumb bob hanging from a taut cord in warm light.

The Experts Who Said It Did Not Affect Them

Everything so far has been about what anchoring is and which version of it survives. This section and the next are about how far the surviving version reaches, because a paradigm that only works on undergraduates guessing at rivers would not be worth this much argument. It works on people paid to know better.

Here is the finding that should worry you more than the wheel of fortune.

In 1987 Gregory Northcraft and Margaret Neale took a real house. They produced a professional information packet on it, the kind an appraiser would normally work from, and they varied one thing: the listed price. Then they took two groups through the property. One group was students. The other was working real estate agents.

Both groups' valuations moved with the listing price. The professionals moved less, but they moved. And when asked afterwards what they had used to reach their figure, the agents overwhelmingly did not mention the listing price [28].

That last part is the important part. It is not that experts are as bad as novices. It is that expertise buys you a partial defence and complete confidence that you did not need one.

The pattern repeats wherever anyone has looked. Anchoring on requested damages in personal injury verdicts, where asking for more gets you more [29]. Anchoring in the courtroom, quantified across the whole literature by a 2021 meta-analysis: twenty-nine studies, ninety-three effect sizes, eight thousand five hundred and forty-nine participants, with an overall standardised effect of 0.58 in studies that included a control group [30].

The same paper reports evidence of possible publication bias, which you should hold in mind alongside the effect size rather than after it. A field that publishes its anchoring successes more readily than its anchoring nulls will produce an inflated average. The authors say so themselves. Reporting the 0.58 without the caveat would be exactly the sin this article is complaining about.

Negotiation is the setting where all of this is least surprising and most exploited. The anchor is the first offer, and whoever makes it usually does better [31].

That result has an awkward implication for the advice everybody gives, which is to let the other side speak first so you can learn what they think. What you learn is real. What you also do is hand them the reference point that both of you will spend the rest of the conversation adjusting around. The same pattern turns up in simulated competitive markets, where an arbitrary starting value shapes where the trading settles [32].

Finance is where the stakes get abstract and the anchors get quiet. There are no opening offers in a stock price. There are previous values, and previous values are anchors.

Studying investors and professionals together, one analysis found that expertise reduces the bias without removing it [33]. That is the Northcraft and Neale result again, in a different suit.

More striking is what happens to forecasts. Consensus forecasts drift toward previously published values rather than being rebuilt from scratch, and because markets trade on consensus, that drift shows up in prices [34]. Analysts' earnings forecasts anchor on prior figures in the same way [35].

Think about what that chain means. A number gets published. Other people's estimates move toward it because it exists. Their estimates become the new consensus, and the consensus moves money. Nowhere in that sequence does anybody have to be careless. Each step is a reasonable person using the best available reference point, and the reference point is partly an echo of the last one.

SettingWho was studiedWhat moved
Property valuationWorking real estate agents and studentsAppraisals tracked an arbitrary listing price
Criminal sentencingExperienced legal professionalsSentence recommendations tracked a dice roll
Civil damagesParticipants judging personal injury awardsAwards tracked the amount requested
Legal decisions overall8549 participants across 29 studiesEffect of 0.58 with control groups, with possible publication bias reported
Emergency medicine108019 patients with heart failureTesting rates tracked a triage note
Financial forecastingAnalysts and investorsForecasts tracked prior consensus values
Peer reviewReviewers of conference submissionsScores tracked previously seen scores

One column is deliberately absent from that table, and it is the one that would say whether any of these professionals knew it was happening. Only one of these studies asked, and in that one they did not. Northcraft and Neale asked directly and the agents named other factors. That combination, a measurable effect plus a confident denial, is what makes anchoring different from an ordinary mistake. An ordinary mistake feels like a mistake at some point. This one never does.

108,019 Patients

The largest piece of real-world evidence in this whole literature is not a laboratory study at all.

Dan Ly, Paul Shekelle and Zirui Song took eight years of national Veterans Affairs data and asked a narrow question. When a patient with congestive heart failure arrives at an emergency department short of breath, does what the triage note says change what the doctor tests for?

Shortness of breath in a heart failure patient has an obvious explanation and a dangerous one. The obvious one is the heart failure. The dangerous one is a pulmonary embolism, a clot in the lungs, which kills people and which you will not find unless you look.

The triage note is written before the physician sees the patient. In 4.1 percent of these visits, that note mentioned the heart failure by name. The study sample was 108,019 patients, with a mean age of 71.9 years [36].

When the note named the heart failure, physicians were less likely to test for pulmonary embolism.

That is anchoring at the scale of a national health system. What the study measured is testing rates rather than outcomes, so the honest statement is that the anchor changed what got looked for, not that it can be counted in missed clots. And it happened to doctors who, like the real estate agents, would tell you a triage note does not decide their differential diagnosis.

What is going on inside the clinician's head is less clear than the population pattern. A randomised crossover experiment with medical residents, using eye tracking alongside diagnostic vignettes, found that fixations on the genuinely discriminating features of a case did not differ significantly between residents who fell for a distracting feature and those who did not [37]. On the eye tracking at least, they looked at the right information as much as anyone else did.

Be careful with that, because it is a null result and this article has already complained about nulls being read as proof of absence. What it licenses is a suspicion rather than a conclusion: if the problem were simply not looking, this is where you would expect to see it, and it did not show up.

This matters for how the problem gets described. If clinicians were failing to notice the evidence, the fix would be attention. If they are noticing it and not updating, the fix is something else entirely. There is also an ongoing methodological argument about whether diagnostic vignette studies can actually separate anchoring from confirmation bias, which are different processes that produce similar-looking errors [38]. If you want the wider picture of how a first impression closes down a diagnosis, our piece on the science behind every missed diagnosis covers the clinical reasoning side in depth, and how the brain defends what it already believes covers the bias that so often follows anchoring rather than causing it.

Older clinical work found the same shape decades earlier, in judgments made under conditions designed to resemble clinical practice [39]. And a 2025 randomised study found that showing an optometrist a patient's previous spectacle prescription shaped what they prescribed next [40].

None of this is a claim about any individual doctor, and none of it is advice about your own care. It is a claim about what happens to human judgment when a number arrives before the evidence does.

What Actually Changes How Strong It Is

If you want to understand anchoring rather than merely fear it, the useful question is not whether it happens but what changes its size. The findings below run from the least intuitive to the one that does the most damage to how this topic is usually described.

The first thing that changes it is precision, and precision does not work the way you would guess. Chris Janiszewski and Dan Uy found that a precise anchor produces less adjustment than a round one [41]. Put crudely, and the round figures here are mine for illustration rather than the study's materials, an asking price that lands on an exact odd number gets negotiated in smaller steps than the same price rounded off.

The proposed reason is that a precise number implies a fine-grained scale, so you adjust in small units rather than large ones. Anyone who has ever wondered why asking prices end in odd digits now has a partial answer.

That one finding is worth more practical attention than most of the advice written about this topic. It is also completely counterintuitive, which is probably why it has not spread.

Plausibility ought to matter more than it does. Wildly extreme anchors still work, though their effectiveness has limits and those limits depend on how the anchor is processed rather than simply on how silly it is [42]. The units matter too: 7300 metres and 7.3 kilometres are the same distance and do not produce the same anchoring [43]. The number itself is doing work, not just the quantity it denotes.

Then there is a set of results that sit awkwardly against everything said earlier, and it would be dishonest to slide past them. Anchors presented below the threshold of conscious perception have been reported to shift judgments [44]. So has the size of a gesture someone makes while speaking [45]. The effect has even been found in mental arithmetic [46], a task with a correct answer the person is perfectly able to compute.

Notice the problem. None of those anchors come with a comparative question attached, which puts them close to the incidental family this article has just finished calling the shakiest part of the literature. You cannot cheer for the subliminal result and dismiss the restaurant result without saying why they differ.

The defensible position is a narrow one. The subliminal and gesture studies are individual findings that have not been through anything like the replication scrutiny the incidental paradigm received, so they are leads rather than load-bearing evidence, and this article treats them that way. If they hold, the process reaches lower than deliberate reasoning. If they go the way incidental anchoring went, the standard paradigm is left carrying the whole phenomenon on its own.

Who you are and what mood you are in changes the size of the effect, though never to zero. Sadness increases susceptibility [47], and in a study of experts specifically, mood and expertise interact rather than one simply overriding the other [48]. Cognitive ability provides some protection but not immunity [49]. So does being certain of the answer, which is why nobody can anchor you on how many days are in a week.

That last point is the one people miss when they describe anchoring as mind control. Anchors do their clearest work where you are genuinely unsure. They do not stop entirely when you are not, which is what makes the virtual reality result and the mental arithmetic result awkward, but certainty is the direction that shrinks them.

There are conditions where the effect vanishes. When the task is framed as an evaluative judgment rather than a numerical estimate, the standard anchoring effect can be eliminated [50]. Careful work on what stimuli are actually necessary for anchoring has narrowed the conditions further [51], and cross-scale versions of the paradigm show the effect depends on semantic relationships that a pure numeric account does not predict [52]. Anchoring is not a physical constant. It has boundaries, and mapping them is most of the current research programme.

And then there is the finding that undermines the most common sentence written about this topic.

A pre-registered set of meta-analyses across four anchoring tasks, using open data from 6,344 participants in ten countries, found strong heterogeneity between cultures. Not noise. Systematic variation, with specific cultural values orientations predicting effect size: Intellectual Autonomy and Egalitarianism were negatively correlated with how strongly people anchored [53].

"A universal feature of human cognition" is the phrase you will find on almost every page ranking for this keyword. Be precise about what the finding does and does not say. Nobody has shown a culture where anchoring is absent. What has been shown is that how big it is varies systematically by culture, which makes "universal" true in the weakest sense and misleading in the sense readers take it.

Identical white stone markers in misty coastal landscape on wet sand.

Where You Meet Anchors Without Noticing

None of the studies above were run on you specifically, so it is fair to ask where any of this touches an ordinary week.

Start with the obvious one. Any negotiation with a number in it has an anchor, and it belongs to whoever speaks first. That is not folk wisdom, it is a measured result: first offers predict where the deal lands [31]. The received advice to never make the first offer is, on this evidence, backwards for most situations. What is true is narrower. If you have no idea what the thing is worth, speaking first exposes you. If you have a reasonable estimate, going first hands you the reference point.

Second, the number does not need to be an offer. It needs to be a number attached to the same question. A previous salary, a former asking price, a competitor's quote, the figure someone mentioned in passing before the meeting started. Each of these arrives without ceremony and then sits underneath the discussion as the value everything is measured against.

Third, precision is a lever and almost nobody uses it deliberately. On the evidence about precise versus round anchors, a specific figure invites small adjustments and a round figure invites large ones [41]. This is why a price that looks calculated tends to be argued with in small increments and a round one invites a round counter-offer.

Fourth, and with the caveat from the previous section still standing: there are reports that an anchor does not even have to be aimed at you, from subliminal presentation [44] to the size of a gesture someone makes while speaking [45]. Those are single findings rather than settled ones and should not be the reason you change anything. The safer point stands without them: most of the anchors in your life are not manipulations, they are accidents of who mentioned what first.

None of that gives you a defence. There is a section further down on what actually works, and most of it is a list of things that were tested and did not. What it gives you is a way to notice. When you find yourself thinking "that seems like a lot" or "that seems reasonable", ask what you are comparing it to, and where that comparison came from. Quite often the answer is a number somebody said out loud eleven minutes ago.

The Part About Studying

Here is the section that almost nothing else written about anchoring contains, and it is the one that should interest you most if you are a student.

Anchors do not only move what you think a house is worth. They move what you think you know.

When you finish studying something and ask yourself whether you have got it, you are making what researchers call a judgment of learning. It is a prediction: how likely am I to recall this later. Those predictions drive everything downstream. They decide what you restudy, how long you keep going, and when you decide you are done.

They are also anchorable, though the picture is more careful than a headline would make it.

This has been tested, and the two most direct tests do not fully agree with each other. The first reports an uninformative anchoring effect on judgments of learning [54]. A later study from the same first author separated the two kinds of anchor and found the split that matters: informative anchoring, where the number carries some relevance, moved binary judgments of learning, while uninformative anchoring did not [55].

Two papers, overlapping authors, different answers on the uninformative case. That is not a scandal, it is what an unsettled question looks like up close, and it is the same fault line running through the rest of this article: anchors with a claim to relevance behave reliably, and anchors that are merely present are where the evidence gets thin.

Then there is study time itself.

Sixty-two Chinese university students studied twenty word pairs under self-paced instructions. The instructions carried an anchor. One group was told that the typical participant had spent fifteen seconds learning each pair. The other was told five seconds. Then everyone studied for as long as they wanted and took a cued recall test. The higher the anchor, the longer they studied, and the longer they studied the better they recalled [56].

So a sentence in the instructions changed how long people worked, and the extra work paid off. The authors call that second half a labor-and-gain effect, and it is the more interesting half: the anchor did not just move a number people reported, it moved what they actually did and what they walked away knowing.

That is a small sample and the finding should be held accordingly. But the design is pointed, because it sits on top of a much older question in the metacognition literature. Thomas Nelson and Jacob Leonesio described the labor-in-vain effect in 1988: students given more study time will use it, and past a point the extra time buys them almost no extra recall [57]. Study time is a decision, not a resource. And decisions can be anchored.

Think about what that means for the way most people study. You sit down without a fixed plan, you work until something tells you that is probably enough, and then you stop. Whatever set your sense of enough was in the room before you started.

The same applies to how hard something feels. Three experiments, with a hundred, eighty-seven and eighty participants respectively, tested whether an anchor shifts how students rate the cognitive load of problem-solving tasks. It did, across tasks of low, moderate and high element interactivity [58].

Your sense of how demanding a task was does not come off a gauge somewhere in your head. It is a judgment, and judgments move.

Put these together and the practical shape is clear enough without any advice attached. The first number you encounter about a task, whether it is how long the typical student takes, or what score the last person got, or how hard the material is supposed to be, becomes the reference point for your own assessment of yourself. That assessment then decides how you spend the next hour.

If you want the wider context for why self-assessment is such an unreliable instrument, our article on the illusion of knowing covers how fluency gets mistaken for mastery, and thinking about your thinking covers the monitoring processes anchors interfere with.

Assessment is anchorable from the other side of the desk as well, and that half of the picture is easier to demonstrate because the judgments get written down.

Researchers have found anchoring effects in how academic papers are assessed, where a prior evaluation shifts the next one [59]. A randomised controlled trial went further and tested peer review itself, and found evidence of reviewer anchoring: a score you have already seen pulls the score you are about to give [60].

That is worth sitting with, because peer review is the mechanism science uses to decide what is true. It runs on the same equipment as everything else.

University rankings show the pattern at institutional scale. Reputation scores anchor on previous reputation scores, which is one reason rankings barely move regardless of what universities actually do [61]. A ranking that is mostly a memory of last year's ranking is not measuring much.

Medical education has taken this seriously enough to propose structured approaches to reducing cognitive bias in assessment, precisely because an examiner's impression of a candidate carries forward [62].

The common thread across all of it is that anchors survive between judgments. They are not confined to the moment. A number from one assessment becomes the reference point for the next, and the chain can run for years before anybody thinks to break it.

So the person grading you is anchored, and so are you, and the anchors are usually numbers nobody thought about very hard.

Antique brass hourglass on dark oak, warm amber sand settled.

What Actually Helps

This is normally where an article gives you five tips. This one is going to tell you what has been tested, including the things that failed, because the failures are more informative.

Being warned does not reliably work. Epley and Gilovich tested forewarning directly and found that it changes behaviour for self-generated anchors, where you can consciously push yourself to adjust further, and does much less for anchors handed to you [63]. This makes sense on the two-process picture. You cannot deliberately adjust away from something that never functioned as a starting point.

Paying people does not reliably work either. Financial incentives for accuracy have been tested since the 1980s and do not remove the effect [64]. The more recent finding that accuracy motivation does shift provided-anchor estimates [16] complicates this rather than reversing it: motivated people do better, and they are still anchored.

Expertise, as you have already seen, is a partial defence sold to its owner as a complete one.

The intervention with the best evidence behind it is considering the opposite. Mussweiler, Strack and Pfeiffer had participants deliberately generate reasons why the true value might be far from the anchor, and the effect was reduced [65].

It works, on this account, because it forces retrieval of anchor-inconsistent information, which is precisely the material the comparative question suppressed. That explanation is selective accessibility, the account this article spent a whole section watching fail its own laboratory test. Say that out loud rather than hoping nobody notices. A theory can be a poor description of what the mind does and still generate an intervention that works, and considering the opposite would keep reducing anchoring even if the reason turned out to be something else entirely. Interventions and mechanisms are scored separately. Later work has tested whether the strategy can be trained rather than merely instructed in the moment [66].

Considering the opposite is not a polite way of saying "be aware of the bias". Awareness is the thing that failed. It is a specific retrieval operation with a specific target, and it takes real effort, which is why it does not survive contact with a busy day unless something in the process forces it.

That is why the more promising work now is structural rather than personal. In medicine, the proposals are simulation training and process changes [67], and one recent study explored using multi-agent conversations between language models to surface alternatives a single reasoner would have skipped [68]. In assessment, the proposal is detailed rubrics rather than vague scales, on the reasoning that a specific criterion anchors the judgment to the work instead of to an impression. Note that this one is a proposal with a mechanism behind it rather than a result, which is exactly the distinction this section is trying to keep.

The honest summary is that individual willpower has a poor record here and design has a better one.

The asymmetry has a cause. Vigilance requires a signal to be vigilant about, and an anchored judgment does not produce one. It feels like an ordinary opinion arriving at ordinary confidence, which is why the real estate agents could answer honestly and still be wrong about themselves.

Procedures do not need the signal. A rubric fixes what you are looking at before you look, and producing your own estimate before you open the email does not ask you to notice anything at all, it just puts the anchor second where it has less to work with. Neither of those is a tested intervention in the way considering the opposite is, and they are offered as the shape the evidence points toward rather than as findings.

Where you have a choice, then, change the order of events rather than trying to be more careful inside them.

Maybe It Is Not a Bug

Everything above treats anchoring as an error. There is a serious argument that this framing is wrong.

Falk Lieder and colleagues modelled anchoring as what happens when a mind with limited time and limited computation has to produce a number [69]. On this account, starting from an available value and adjusting until the expected gain from further adjustment falls below its cost is not a malfunction. It is the correct policy for a resource-limited system. The model predicts insufficient adjustment, and it predicts that adjustment gets less sufficient under time pressure and cognitive load, which is what the data show.

If that is right, then anchoring is not a flaw bolted onto reasoning. It is what reasoning looks like when reasoning is expensive.

This is not a fringe position, and it changes what a mitigation is supposed to do. You are not repairing a broken component. You are deciding, case by case, which judgments are worth more computation than your brain would spend on them by default. Most are not worth it, and the few that are tend to be the ones with a number and a signature attached.

The strongest brain evidence comes from a place you would not look for it. Guessing what another person thinks means starting from what you think and adjusting, which is structurally the same operation, and Diana Tamir and Jason Mitchell found the neural signature of that adjustment while people did exactly that [70], then extended it to social inference generally [71]. So the machinery is not specific to prices and numbers. A separate group stimulated the right dorsolateral prefrontal cortex and reported changes in anchoring [72], which is one study and should be read as a lead rather than a location. Domain knowledge changes judgment without switching any of this off, which is the subject of our piece on how expertise is actually built.

The Machines Do It Too

A short section, because the literature is young and moving.

Large language models show anchoring. In a forced-choice diagnostic task, models exhibited greater diagnostic anchoring than physicians did [73]. Prompts that contain a cognitive bias reduce model accuracy on radiology board-style questions [74], and cognitive biases have been evaluated in AI-assisted mammography reading [75].

The more interesting result is what happens when a person and a model work together. Research on AI-assisted decision-making found that cognitive biases shape how people use a model's suggestion, so an accurate recommendation delivered first can function as an anchor [76]. Related work on algorithmic fairness found that anchoring limits what information transparency can achieve, because showing people more does not stop the first number from doing its work [77].

There is a second reason to care, which has nothing to do with psychology. If a model anchors, then the order in which you feed it information changes what it tells you, and most people using these systems assume the opposite. They assume the model reads everything and weighs it. Give it your hypothesis first and you may be getting a version of your own hypothesis back with citations attached.

Do not over-read this. A model reproducing an anchoring pattern does not prove anything about human mechanism. Language models learn from text written by anchored humans, so the finding is at least as much about the training data as about cognition. What it does establish is practical: if you are using a model to check your reasoning, the order in which you give it information matters, and giving it your current hypothesis first is a way of getting it back.

Smooth grey pebble on still dark water with ripples.

What We Actually Know

The state of play, stated plainly, with the disputes left in.

Settled well enough to state without hedging: the standard two-step anchoring paradigm produces reliable effects, and has for fifty years. Irrelevant anchors work. Experts are affected and tend not to notice. Anchor precision changes how far people move. Awareness and incentives are not reliable fixes. Considering the opposite has the best supporting evidence of anything tested.

Unsettled, and honestly so:

Whether selective accessibility explains anchoring is open. Strack, Mussweiler and colleagues built the account and supported it across a decade of work. Harris and colleagues could not find its signature in five properly powered experiments, and other results point to mutual assimilation that the account does not predict. The theory is not refuted. Its best-known test is.

Whether adjustment is the mechanism for self-generated anchors is mostly supported and partly complicated, with Epley and Gilovich on one side and Simmons and colleagues showing that the clean split leaks.

Whether incidental environmental anchoring is a real phenomenon is genuinely in doubt. Critcher and Gilovich reported it, Many Labs 2 failed to replicate the key study, and Shanks and colleagues found nothing across three attempts while finding standard anchoring in the same work.

How large the effect is depends on where you look, and the legal meta-analysis that gives the best single number also reports possible publication bias in its own literature.

Whether it is universal is answered: not uniformly. Effect sizes vary substantially across cultures in a way that tracks cultural values.

And whether it is a defect at all is a live philosophical question with a formal model behind it.

That is what fifty years of work on one of the most studied effects in psychology actually produced. Not a clean law. A reliable phenomenon, a contested explanation, a discredited extension, and a moving boundary. It is a less satisfying story than the one you have read elsewhere, and it is the one that is true.

The practical upshot is smaller than the usual advice and more useful. You cannot stop numbers from reaching you first. What you can occasionally do, on the decisions that are worth the effort, is produce your own estimate before you look at anyone else's, and then deliberately spend two minutes arguing for a number far from the one you were given. It is a narrow tool, and it is not the only thing with any support. Forewarning does help when the anchor is one you generated, and caring about accuracy does move the estimate. Considering the opposite is simply the one with the best evidence and the clearest account of why it should work at all. If you are interested in the wider family of judgment errors, the statistical artifact hiding inside a famous chart is another case where the popular version and the evidence have drifted a long way apart.

Frequently Asked Questions

What is the anchoring effect?

The anchoring effect is the tendency for the first number you encounter to pull your later judgments toward it, even when that number is irrelevant to the question. The founding demonstration spun a wheel of fortune in front of participants and then asked them to estimate the percentage of African countries in the United Nations. People who saw a higher number gave higher estimates. The number carried no information about the United Nations at all. It moved them anyway. The standard laboratory version asks you to compare a quantity against a number first and then estimate it, and that version has replicated reliably for fifty years across pricing, forecasting, medicine, law and negotiation.

What is the difference between anchoring bias and availability bias?

They are different failure modes of the same underlying process, which is judgment under uncertainty running on shortcuts. Anchoring is about a starting point: a number arrives, and everything you produce afterwards is measured relative to it. Availability is about what comes to mind: you judge how likely something is by how easily examples surface, so vivid or recent events feel more common than they are. Both were described in the same 1974 paper by Tversky and Kahneman. The practical difference is that anchoring needs a number and availability does not, and that anchoring can be introduced deliberately by whoever speaks first.

Why does the anchoring effect happen?

This is the part most articles get wrong by sounding certain. There are two leading explanations and neither has won. Selective accessibility says the anchor works as a search term: testing it as a candidate answer activates knowledge that fits it, and you then answer from that biased sample. Anchoring and adjustment says you start at the anchor and move away from it, stopping too early. The evidence currently suggests both are partly right for different kinds of anchor, with adjustment doing better for numbers you generate yourself and accessibility doing better for numbers handed to you. What has changed recently is that five properly powered experiments failed to find the standard laboratory signature of selective accessibility, while finding the anchoring effects themselves. The mechanism is genuinely unsettled.

Is the anchoring effect real or did it fail to replicate?

Both statements are true about different things, which is why the topic is so confusing. The core effect is real and keeps replicating, including in a 2026 study where people judged how high up they were standing in a virtual environment with the evidence in plain view. What failed is more specific. Many Labs 2 did not replicate a key incidental anchoring study, and three later studies found no statistically significant incidental environmental anchoring while finding strong standard anchoring in the same work. Incidental anchoring is the version where a number merely present in the environment supposedly changes what you would pay, and it is also the version consumer articles like most. The laboratory paradigm survived the replication era. That one branch of it did not.

Does knowing about the anchoring effect help you avoid it?

Not by itself. Forewarning has been tested and it helps mainly with anchors you generate yourself, where you can consciously push the adjustment further, and does much less for numbers handed to you. Financial incentives for accuracy have been tested since the 1980s and do not remove the effect either. Expertise reduces it without removing it, and in the one study that asked directly, the professionals did not name the anchor as something they had used. The one strategy with real support is considering the opposite: deliberately generating reasons the true value might be far from the anchor, which forces you to retrieve information the comparison suppressed. It takes effort, so it works best when it is built into a process rather than left to willpower.