Introduction
You have seen the chart. A curve climbs steeply to a peak on the left, labelled something like Mount Stupid. It falls into a trough called the Valley of Despair. Then it rises again, slowly, along a gentle slope toward the Plateau of Sustainability. Somewhere on that image are the words Dunning-Kruger.
It is one of the most shared pictures in popular psychology. It has been used in management training, in medical education, in arguments online, and in at least one talk you have probably sat through.
It is not in the paper.
Not a version of it. Not a simplified redraw of it. The 1999 study that gave the effect its name contains four figures, and every one of them plots something else entirely. There is no peak. There is no valley. Both lines on those figures rise from left to right, quietly, in parallel, like two people walking uphill at slightly different speeds.
That is the small problem. The large problem is what happens when you look at the way those figures were drawn, because the method has a property that almost nobody mentions. Run it on numbers generated at random, numbers with no human being anywhere near them, and it produces the same shape.
This article is about what the original researchers measured, what their critics found when they measured it differently, and what is left standing in 2026 after twenty-seven years of argument. The answer is not that a famous finding was exposed as fake. It is stranger and more useful than that. The instrument was broken, the explanation was wrong, and something real is still there underneath, wearing a different shape than the one you were sold.

What Kruger And Dunning Actually Did
In 1999 two psychologists at Cornell University, Justin Kruger and David Dunning, published a paper in the Journal of Personality and Social Psychology with a title that did most of the work of making it famous: "Unskilled and unaware of it" [1].
The design was simple. Give people a test. Before or after the test, ask them to estimate how well they did compared with everyone else, on a scale from 0 to 100. Then compare the estimate with the result.
They ran it four times. The first study used humour, which sounds soft until you see how they handled it. They took thirty jokes, had eight professional comedians rate each one, and used the comedians' consensus as the standard. Sixty-five undergraduates then rated the same jokes and estimated their own standing. The second study used a twenty-question logical reasoning test taken from a law school admissions preparation guide, with forty-five undergraduates. The third used grammar, judged against the conventions of American Standard Written English, with eighty-four undergraduates. The fourth returned to logical reasoning with one hundred and forty undergraduates and added something the first three did not have, which we will get to.
Add those up. Three hundred and thirty-four students at one selective private university in upstate New York in the 1990s.
The Numbers That Made It Famous
The finding was consistent across all four. People who scored badly thought they had done fine.
In the humour study, the whole sample put their ability to spot what is funny at the sixty-sixth percentile on average. That is already sixteen points above where a group average has to sit, since by definition the average person is at the fiftieth percentile. The bottom quarter of scorers landed around the twelfth percentile in reality. They estimated themselves at the fifty-eighth. That is an overestimate of forty-six percentile points.
The logic study repeated it. Bottom-quarter scorers put their general logical reasoning ability at the sixty-eighth percentile and their test score at the sixty-second, against an actual standing near the twelfth. The grammar study found the same shape, with the whole sample estimating their grammar ability at the seventy-first percentile and their test performance at the sixty-eighth.
Those numbers are the reason the effect entered the language. A person in the bottom eighth of a group, sincerely reporting that they are in the top half, is a memorable thing.
Look at the size of that foundation for a moment. It is four small studies. Not one of them exceeds one hundred and forty people, and all of them drew from the same student population. That is a perfectly normal size for a social psychology experiment in 1999, and it was never presented as anything else. What happened afterwards was not the authors' doing. A modest set of campus experiments became, in public, a law of human nature that explains your colleagues, your relatives and every stranger who has ever disagreed with you online.
The Half Of The Finding Nobody Quotes
Here is the part that fell out of the story on its way to the internet.
The top quarter of scorers were wrong too. They underestimated themselves. In every one of the four studies, the people who actually performed best placed their relative standing lower than it really was.
That does not fit the meme at all. The popular version says confidence and competence run in opposite directions, so the fools are certain and the experts are humble. What the data showed is that everyone's estimate was squashed toward the middle. Low scorers guessed too high. High scorers guessed too low. Both groups were closer to average in their own minds than they were on the test.
Kruger and Dunning had an explanation for the second half, and it was not the same as their explanation for the first. Top performers, they argued, were accurate about themselves and wrong about everyone else. Having found the task easy, they assumed it was easy for others too. That is a false consensus problem, not a metacognition problem. The single label covering both halves was always doing more work than it could carry.
Keep this in mind, because it is the earliest sign that the picture is not what it seems. A pattern where both ends of the scale are pulled toward the middle has a name in statistics, and we will come back to it.
The Explanation They Offered
Why would poor performers get it so wrong?
The answer Kruger and Dunning gave is the one that made the effect feel profound. They proposed that the skills you need to do something well are largely the same skills you need to recognise that you have done it well. If your grammar is poor, you cannot spot the errors in your own writing, because spotting them is the very ability you lack. Incompetence, on this account, is self-concealing.
Dunning later called this the dual burden [2]. You fail at the task, and the failure hides itself. He and his colleagues developed the argument across a long review of self-assessment in the early 2000s [3] and defended it against early challenges with new experiments [4].
It is a genuinely good idea. It is testable, it makes predictions, and it explains something people recognise from life. It also happens to be the part of the theory that has taken the heaviest damage, which we will get to in its proper place.
Study 4 was where they tried to prove the mechanism directly. They gave a group of students a short training session in logical reasoning, then asked them to re-estimate their earlier performance. Poor performers who had received the training became more accurate about how badly they had done. Teach someone the skill and they gain the ability to see they lacked it. As a demonstration of the dual burden, that is about as clean as psychology gets.
The Training Study, Reconsidered
That fourth study is the strongest card in the original hand, and it is worth understanding why the critics were not persuaded by it either.
The design asks a group of people to estimate their performance, teaches them something, and asks again. The people whose estimates improve most are, by selection, the people whose first estimates were furthest from their scores. That is a group defined by having an extreme value on the first measurement.
You already know what happens to extreme groups on a second measurement. They move toward the middle. Some of that movement is real learning. Some of it is the same drift that happens to any extreme group measured twice, with no teaching involved at all. A study that lacks an untrained control group measured twice cannot separate the two.
Kruger and Dunning were aware of this general class of objection and argued in their 2002 reply that regression could not be doing all the work. The disagreement was never about whether the phenomenon exists in the arithmetic. It was about proportion. How much of the gap is the graph, and how much of it is the person.
For nearly two decades that question did not have a clean answer, because answering it required methods that were not standard in the field yet. When those methods arrived, they did not settle it in the direction the original authors expected.
There is a broader lesson sitting in here about how psychology arguments go. Both sides had the same data for twenty years. Both sides were arguing in good faith. What changed the picture was not new evidence about human nature. It was a better understanding of what the old evidence was capable of showing.
The First Objection Arrived Three Years Later
In 2002, Joachim Krueger and Ross Mueller published a paper in the same journal with a blunt question in the title: "Unskilled, unaware, or both?" [5].
Their argument had two parts, and neither one required doubting a single number in the original.
The first part was the better-than-average effect. On tasks that people believe are easy and common, most people rate themselves above average. This is one of the oldest and most reliable findings in social psychology, and Kruger himself had published on its mirror image the same year the famous paper appeared, showing that on tasks people believe are hard the pattern flips and most people rate themselves below average [6]. If nearly everyone claims to be above average, then the people who are genuinely at the bottom will show the biggest gap between claim and reality. Not because they are blind to their failings. Because they are at the bottom, and the general upward bias has further to travel from there.
The second part was regression to the mean, and this is where it gets uncomfortable.
That second part deserves its own section, because it does more work in this argument than anything else and gets less attention in public accounts of it than anything else.
Regression To The Mean In Plain Language
Suppose you measure something with any amount of error in it. Test scores have error. Self-ratings have error. Everything measured about a person on a Tuesday has error.
Now take the people who scored lowest. That group contains two kinds of people mixed together: those who really are the weakest, and those who are ordinary but had a bad day, misread a question, or guessed wrong twice. When you measure that group again on anything else, the bad-day people drift back toward their true level, which is higher. The group average moves up.
The same thing happens at the top, in reverse. The highest scorers include genuinely strong performers plus lucky ones, and the lucky ones drift down.
Now put a self-rating next to it. A self-rating is a second measurement, taken with different error, of roughly the same underlying thing. So the bottom group's self-ratings will sit above their test scores, and the top group's self-ratings will sit below their test scores, for the same reason a second exam would.
You have just produced the entire Dunning-Kruger picture. Both ends pulled toward the middle. Low scorers apparently overconfident, high scorers apparently modest. And you did it without any claim about metacognition, self-knowledge, or the psychology of incompetence.
Kruger and Dunning replied in the same issue that year. Statistical factors could not account for the whole pattern, they argued, and the training study in their fourth experiment pointed to something genuinely metacognitive underneath [7]. The argument has run ever since without closing.
That is the objection in one paragraph. The response to it, for years, was that regression can explain some of the gap but surely not a forty-six point one. Then people started checking exactly how much it could explain.
The Score Is On Both Axes
Here is the mechanical heart of the problem, and it is worth slowing down for, because almost no popular account of this argument explains it.
Look again at how the classic chart is built. You sort people into four groups by their test score. Then, for each group, you plot two things: their average test score and their average self-rating. The x-axis is the test score. One of the two lines on the y-axis is also the test score.
The test score is on both axes.
This is not a subtle statistical concern. It is the same variable appearing twice in the same picture. Once you sort people by a number, that number stops being a free variable and becomes the thing that defines the group. Anything you plot against it inherits the sorting.
It gets worse when the chart plots the gap. Many papers and almost every popular redraw show the difference between self-rating and score, plotted against the score. Call the self-rating y and the score x, and you are plotting y minus x against x. The x on the left of that expression and the x on the right are the same numbers. Correlate a variable with itself and you get a relationship, guaranteed, before any data about people enters the calculation.
Edward Nuhfer and his colleagues named these conventions precisely in a 2016 paper in the journal Numeracy, listing the offenders: scatterplots of y minus x against x, column graphs of y minus x against x aggregated into quantiles, and line charts of data aggregated into quantiles [8]. Their conclusion was that these graphical conventions "introduce artifacts that invite misinterpretation". A follow-up paper the next year traced the consequence: random noise plus one graphical habit had been enough to send a field's explanations of self-assessment data in the wrong direction [9].
There is a second constraint stacked on top of the first, and it is purely arithmetic. Percentile scales have ends. A person who genuinely scores at the third percentile cannot overestimate themselves downward. There is almost no room below them. The only direction their error can point is up. A person at the ninety-seventh percentile has the mirror problem: their error can only point down. Even if everyone in the study made random errors of exactly the same size, the errors at the bottom would all be positive and the errors at the top would all be negative, and the chart would show low performers overestimating and high performers underestimating.
Jan Magnus and Anatoly Peresetsky built precisely this into a formal model in 2022, taking the random boundary constraints into account explicitly and requiring no psychological ingredient whatsoever. Their paper reports that the model fits the data almost perfectly [10]. A model with no psychology in it reproducing a psychological finding is not a small result.
The Random Number Test
If the objection above is right, there is an obvious way to check it. Skip the humans. Generate numbers at random, pretend one column is a test score and the other is a self-rating, and draw the chart.
Nuhfer and colleagues did this [8]. The familiar picture appeared. Low scorers overestimating, high scorers underestimating, the two lines converging toward the right exactly as they do in the published literature. Nobody in the dataset had any metacognitive ability at all, because nobody in the dataset existed.
Sit with that for a second, because it is the single most useful fact in this whole argument, and it is routinely misunderstood in both directions.
It does not prove the Dunning-Kruger effect is fake. That is the overclaim, and it is everywhere. What it proves is that the chart cannot tell you whether the effect is real, because the chart looks the same either way. A test that returns positive when the thing is present and also when it is absent is not a test. It is a decoration.

A Worked Example You Can Follow
If the argument still feels abstract, here is the whole thing in numbers small enough to hold in your head. Nothing below is data from a study. It is an illustration of what the arithmetic does on its own.
Imagine one hundred people take a fifty-question quiz. Imagine, for the sake of the example, that everybody guesses completely at random, so nobody knows anything and nobody has any insight into their own performance whatsoever. Scores will still vary. Some people will get thirty by luck. Some will get twenty.
Now ask everyone the same question: what percentile do you think you landed in? Since nobody has any real information, imagine their answers cluster loosely around the middle, somewhere in the forties and fifties, with a bit of spread.
Now build the classic chart. Sort the hundred people into four groups by their quiz score. Plot each group's average score and each group's average self-estimate.
The bottom group scored near the tenth percentile because that is how they were defined. Their self-estimates still sit around the middle, because their self-estimates carry no information about anything. So the bottom group appears to overestimate itself by roughly forty points.
The top group scored near the ninetieth percentile, again by definition. Their self-estimates also sit around the middle. So the top group appears to underestimate itself by roughly forty points.
Draw the two lines. One rises steeply, because it is the sorting variable. One stays flat, because it is noise. They cross in the middle and the gap is widest at the ends.
You have just reproduced the entire published pattern, including the detail that gets treated as the most psychologically interesting part, the way high performers are modest and low performers are not. There were no high performers. There was no modesty. There was a sorting operation and a flat line.
Real people are not this extreme, obviously. Real self-estimates carry some genuine information, which is why the self-rating line in published studies slopes upward instead of lying flat. That slope is the real signal, and it is the thing worth measuring.
But notice what the classic chart does with that signal. It buries it. The chart's most eye-catching feature, the widening gap at the bottom, is produced by the sorting whether or not any signal exists. The feature that actually tells you something, the slope of the self-rating line, is the part nobody points at.
This is why the modern methods look the way they do. One asks whether the spread of self-assessment errors widens as ability drops. The other asks whether the line bends rather than running straight. Both questions go at the signal directly and leave the sorting alone.
Difficulty Flips The Whole Thing
While the statistical argument was building, a separate line of work was attacking the finding from the side.
Katherine Burson, Richard Larrick and Joshua Klayman published a paper in 2006 whose title tells you where it is going: "Skilled or unskilled, but still unaware of it" [11]. Their point was that perceived task difficulty drives the whole pattern, and that you can reverse it at will.
Give people an easy task and everyone thinks they did well, so the worst performers show the biggest overestimation. Give people a hard task and everyone thinks they did badly, so the best performers show the biggest underestimation and the pattern turns upside down. The effect, on this reading, is not about the relationship between skill and self-insight. It is about how hard the task felt to everyone in the room.
Marian Krajc and Andreas Ortmann made a related argument from economics two years later, offering an alternative account of why the unskilled might look so unaware without any deficit in self-knowledge [12].
Around the same time, Don Moore and Paul Healy published a reorganisation of the whole overconfidence literature in Psychological Review, separating three things that had been getting mixed together: thinking you performed better than you did, thinking you rank higher than you do, and being too certain that your estimate is correct [13]. Those three come apart. They can even point in opposite directions in the same study. A lot of the confusion in this field comes from articles that measure one and conclude about another.
Not every challenge succeeded. Thomas Schlösser and colleagues tested a specific alternative explanation, the idea that self-assessments are a rational signal extraction problem, and reported that it did not account for the data as well as its proponents hoped [14]. That result cuts toward the original account. The record here genuinely goes both ways, which is why anyone telling you this is settled is selling something.
What Happens When You Measure It Properly
By around 2019 the argument had a clear shape. Critics said the quartile chart manufactures the pattern. Defenders said the pattern is bigger than the artifact can explain. The way to settle it is to stop using the chart and use methods that do not have the flaw.
Three groups did exactly that, and their results are the backbone of the modern position.
Gilles Gignac and Marcin Zajenkowski took 929 members of the general community, measured their intelligence with the Advanced Raven's Progressive Matrices, and asked them to self-assess. Then, instead of sorting into quartiles, they used two techniques that test the Dunning-Kruger prediction directly. The Glejser test asks whether the spread of self-assessment errors gets wider at lower ability, which is what the hypothesis requires. Quadratic regression asks whether the relationship between real and perceived ability bends, which is what the hypothesis also requires. They found no significant heteroscedasticity and an association that was essentially entirely linear. Their title said it plainly: the effect is mostly a statistical artefact [15].
Robert McIntosh and colleagues went after the explanation rather than the pattern. In a 2019 paper they clarified what role metacognition could actually be playing [16], then followed it with a Registered Report, meaning the analysis plan was locked and reviewed before the data existed. They recruited 159 people and obtained 151 valid datasets on a matrix reasoning task, using modern measures that separate three different things a self-rating can reflect: sensitivity, efficiency and bias.
The results are worth stating carefully.
Metacognitive sensitivity, meaning how much real information a person's confidence judgements actually contain, tracked performance closely. Poor performers had less information to work with. Metacognitive efficiency, meaning the quality of the metacognitive machinery itself once you account for how much information was available, was unrelated to performance. And metacognitive bias, the general tendency to be confident or unconfident, was positively associated with performance, meaning poor performers were less confident than good performers, not more.
Then the crucial finding. None of those metacognitive factors caused the classic pattern. It was driven overwhelmingly by the performance scores themselves. The authors concluded that the classic effect is a statistical regression artifact that tells us very little about metacognition [17].
That is a direct empirical refutation of the dual burden, using the method the dual burden's own logic demands. Poor performers were not blind to their failings. They knew. They simply did not adjust far enough.
The sequence above is the whole dispute compressed. One number does two jobs. It defines the groups and it grades the members of those groups. Everything downstream of that double duty inherits the problem.
The Same Data Says Yes And No
The cleanest demonstration that the method decides the answer came in 2024, and it did not come from either camp's usual suspects.
Izabela Lebuda, Gabriela Hofer, Christian Rominger and Mathias Benedek studied the effect in creative thinking, across two studies with 425 and 317 participants. They ran the classical quartile analysis. Then they ran the Glejser test and quadratic regression on the same data.
The classical analysis supported the Dunning-Kruger effect. The modern analyses did not [18].
Same people. Same answers. Same self-ratings. Two verdicts, decided entirely by which statistical tool came out of the drawer.
If you take one thing from this article, take that. The argument is not about whether people are honest about themselves. It is about whether a particular chart is capable of answering the question, and the answer is that it is not.
Gignac and Zajenkowski responded to a critic in 2023 with a paper titled, with some weariness, "Still no Dunning-Kruger effect" [19]. Gignac followed it in 2024 with a reassessment concluding that the effect has negligible influence and applies at most to a limited segment of the population [20]. Note what that is not. It is not zero. It is a much smaller claim than the original, defended by the person who did the most to shrink it.
The other direction has evidence too. Curtis Dunkel, Joseph Nedelec and Dimitri van der Linden published a reevaluation and replication in the same journal that reached more sympathetic conclusions [21], and the same group later used a twin design to ask whether the accuracy of a person's self-judged intelligence is heritable at all [22].
So Is Anything Left
Yes. And this is where most of the debunking coverage stops early and gets it wrong.
The first thing left is a genuine defence of a real effect, published in Nature Human Behaviour in 2021. Rachel Jansen, Anna Rafferty and Thomas Griffiths built a rational model of the whole situation and argued that the pattern is consistent with low performers being insensitive to the evidence available to them about their own performance [23]. Their account does not need incompetence to be self-concealing in the mystical way the dual burden suggests. It needs only that weak performers receive weaker signals and update on them too little.
The second thing left is more striking, and it comes from a domain built for this question.
Consider tournament chess. Every rated player carries a number that is public, precise, updated continuously, and derived from thousands of objective outcomes against opponents whose numbers are equally public. If accurate feedback cures overconfidence, chess should be the cure.
Patrick Heck and colleagues surveyed 3,388 rated players from twenty-two countries, aged five to eighty-eight, with an average of 18.8 years of tournament experience. On average, players asserted an ability 89 rating points above what their actual rating indicated, which corresponds to expecting to beat an equally rated opponent roughly two games to one. A year later, only 11.3 percent had reached the rating they had asserted. Low-rated players overestimated the most. Top-rated players were calibrated [24].
Read that shape carefully, because it is the original finding's shape, arriving through a completely different door. It was not produced by sorting people into quartiles on a twenty-question quiz. It was produced by comparing a stated belief against a number that already existed, measured over years, for thousands of people.
The third thing left is a correction to who is most overconfident. Carmen Sanchez and David Dunning ran a series of studies on people learning something new and found that overconfidence does not peak among the completely unskilled. It peaks among beginners, once a little knowledge has arrived and before enough has arrived to reveal the size of the field [25]. Dunning returned to the theme with a review arguing that intermediate knowledge, not zero knowledge, is what predicts overconfidence [26].
That finding is quietly devastating to the meme while rescuing something from it. The person most at risk of an inflated self-assessment is not the one who knows nothing. It is the one who has just finished the introductory course.
And underneath all of it sits the calmest result in the literature. Ethan Zell and Zlatan Krizan pooled findings across a very large body of self-assessment research and concluded that people have moderate insight into their own abilities [27]. Not excellent. Not absent. Moderate. That is the honest headline, and it will never go viral.
Where The Effect Keeps Getting Reported
One reason this argument matters is the sheer volume of research that has applied the original method to new fields without inheriting the criticism of it.
Medicine produces a steady stream of it, and you can see why. Clinical training runs on self-assessment. Trainees are constantly asked to judge whether they are ready for the next responsibility, and the cost of a bad judgement is not an embarrassing dinner party.
Researchers have reported the pattern in medical trainees generally [28] and in first semester medical students specifically [29]. A study of emergency medicine residents took a sharper approach, comparing what residents predicted they would score on an in-training examination against what they actually scored [30]. That design has a virtue worth noticing. The prediction and the outcome are separate events, recorded at different times.
The response inside medical education has been practical rather than theoretical. A broad review of how physicians maintain expertise examined the strengths and weaknesses of self-assessment as a tool for that job [31]. A systematic review asked whether showing doctors video of their own performance makes their self-assessments more accurate [32]. Both treat self-assessment as a skill that can be improved, which is a more useful framing than treating it as a fixed human flaw.
The other health professions ask the same question for the same reason: competence here is licensed, so somebody has to certify it, and self-report is the cheapest instrument available. Studies have compared health professions students' self-assessments against objective measures of their competence [33], and nurse educators have tested whether the format of the training changes how overconfident the trainee ends up [34].
Patients get studied too, which is where it starts to matter for people who never volunteered for research. One study looked at overconfidence in managing your own health concerns alongside how much health literacy you actually have [35]. Another examined how older adults' subjective sense of their own health related to whether they kept smoking [36]. The worry in both is not that someone feels clever. It is that a confident misjudgement changes what a person does next.
Face perception produced the most instructive episode, because the whole argument played out in miniature. Researchers reported Dunning-Kruger effects in how well people judge their own ability to match faces [37]. Others then wrote to the journal objecting specifically to the use of performance quartiles in that work [38]. The matter was later revisited in a Registered Report designed to investigate individual differences in face matching and self-insight more carefully [39].
That sequence is what a healthy field looks like. Claim, methodological objection, preregistered rematch.
Reasoning research has its own thread. Overconfidence has been tied to the kind of intuitive errors people make on the cognitive reflection test, where a question has an obvious wrong answer that feels right [40]. Later work used the same instrument to look at how confident people are in those intuitive answers [41]. A 2024 study went one level up, examining second-order judgements, meaning what people think about their own estimates, and found these increase with miscalibration among low performers [42].
Money produced one of the more useful null results. When researchers compared what people actually know about finance against what they think they know, they reported a failure to observe the Dunning-Kruger effect at all [43]. That is a domain with real consequences, a large sample and a clear answer, and it is the kind of result that rarely travels.
Related work has asked how accurately people estimate different aspects of their own intelligence [44] and whether cognitive ability predicts how badly calibrated a person's financial expectations turn out to be [45].
Then there is the political and informational cluster, which is where the term does the most public work and needs the most care.
Studies have reported the pattern in the endorsement of anti-vaccine policy attitudes [46] and in anti-consensus views on contested scientific questions [47]. Both are studies of people who are confident about a technical question and wrong about it.
A related strand connects overconfidence to conspiracy endorsement and distrust of science [48] and to susceptibility to false news [49]. Read these carefully before repeating them. Finding that people who score low on a knowledge test are confident anyway is a different claim from finding that confidence caused the belief.
Political knowledge has been examined directly, first through the lens of political sophistication [50] and more recently in a 2025 reassessment [51]. Researchers have even studied it among people who believe the earth is flat [52].
A related question is whether people can tell good information from bad. Overconfidence in the ability to spot misinformation has been studied on its own terms [53], alongside work on why people fall for it in the first place [54] and on how metacognitive awareness relates to detecting confident nonsense [55].
The newest wave has turned the method on machines, and the results are odd enough to be worth a moment. Large language models have been reported to show Dunning-Kruger-like behaviour when fact-checking across languages, being most confident where they are weakest [56]. A 2026 study of people working alongside artificial intelligence reported the same split its title describes, with performance and self-knowledge moving apart rather than together [57]. Whatever the classic chart measures, it is apparently not confined to humans.
The workplace gets its share. One 2025 study reported reduced susceptibility to the effect among autistic employees [58], which is an interesting result to sit with given how often the popular version of this idea is aimed at people who process things differently. Management research has looked at the illusion of competence among managers [59].
You should hold all of that lightly, and the reason is now familiar. Many of these studies used the quartile method, which means they inherit every problem described above. That does not make their subject matter uninteresting or their authors careless. Some of them are careful about it. It means the correct reading of a headline like "the Dunning-Kruger effect found in X" is usually "the standard chart was drawn for X and looked the way the standard chart looks".
How To Read A Study That Reports This Effect
Since this pattern is reported somewhere almost every month, here is a short guide to reading those reports without being misled in either direction.
Ask first what was compared. If a study sorted people into groups by their score and then compared those groups' self-ratings against the same score, you are looking at the classic design and everything in this article applies to it. If it correlated self-rating with performance across everyone without sorting, or used a test of whether error spreads at low ability, or used a pre-existing external measure like a rating built from years of results, then it is doing something the criticism does not automatically cover.
Ask second whether the sample can bear the conclusion. A great deal of this literature runs on undergraduates in a single course, which is fine for a first look and thin for a claim about humanity. Sample sizes in the low hundreds are the norm rather than the exception.
Ask third what the effect is being used to explain. There is a large difference between a paper reporting that self-assessment accuracy is imperfect in a training context, which is well supported and practically useful, and a commentary using the phrase to explain why a group of people the author dislikes holds the opinions they hold. The first is research. The second is the meme wearing a lab coat.
Ask fourth whether the authors mention the dispute at all. Papers published after about 2020 that use the quartile method without acknowledging that the method is contested are, at minimum, behind on their reading. Plenty of recent work does acknowledge it and proceeds carefully, which is the correct response to a contested method.
None of this means the field is broken. It means a specific tool got adopted faster than its limits were understood, which happens in every discipline. The useful posture is neither to believe every report nor to dismiss the whole subject. It is to ask what the picture was drawn from.

Why Feedback Does Not Fix It
If the effect were purely an artifact of drawing, feedback would be beside the point. If it were purely a metacognitive deficit, feedback should cure it. Neither prediction holds cleanly, which is itself informative.
Jennifer Osterhage tracked calibration through a course that gave students repeated practice tests with feedback. Miscalibration persisted for low and high achievers alike [60]. The students were not deprived of information about their performance. They received it repeatedly and stayed miscalibrated anyway.
The chess players make the same point on a much larger scale [24]. Two decades of public numerical feedback is about as much correction as any human being receives about any skill, and it left a gap of 89 rating points.
So the interesting question is not whether feedback exists. It is what kind of feedback would have to arrive, and in what form, before a belief moves.
Some of the attempts are worth knowing about. Researchers have tested whether paying people to be accurate improves how well students calibrate their expectations [61]. Others built a tool that let university students watch their own learning as it happened, on the theory that better information about yourself should produce better estimates of yourself [62]. One economics instructor simply taught students about overconfidence directly, to see whether knowing about a trap is enough to climb out of it [63].
Nobody has reported a clean fix. That is the honest summary, and it is worth stating plainly rather than dressing up.
The economists have made the sharpest contribution here, because they went after the measurement rather than the person. One study estimated how much of the apparent relationship between skill and overconfidence survives once you correct for measurement error in the skill estimate itself [64]. That is regression to the mean again, wearing an economist's suit. Related work has tested overconfidence in simple real-effort tasks where the output is easy to count [65], and a 2025 paper rebuilt the measure entirely, using composite indices rather than a single quiz and reporting that men and women do not come out of the analysis the same way [66].
Two cautions belong at the end of this, and they cut in opposite directions.
The first is that some of what looks like poor self-knowledge may be poor questionnaires. Kit Double published an analysis in 2025 arguing that survey measures of metacognitive monitoring are frequently invalid [67]. If the instrument is noisy, the person looks confused.
The second is that self-monitoring is not one fixed human trait that either works or does not. Work on how people regulate metacognitive effort across cultures suggests a good deal of it is learned and situational [68]. Intellectual humility has been studied along similar lines, as something a person can have more or less of rather than a switch [69]. If that is right, then the reason feedback so often fails is not that people are incorrigible. It is that we keep delivering it in the one format that flatters the existing belief: a number, at the end, with no account of what produced it.

What This Means For Judging Your Own Skill
There is no self-help section coming. But the research does support a few plain statements, and they are more useful than the meme ever was.
Your confidence is a poor instrument for measuring your knowledge, and it is poor in a specific way. It is not that confidence is random. It is that confidence tracks how familiar something feels, and familiarity is not the same as competence. This is the same trap described in the illusion of knowing, where re-reading a page until it feels smooth produces a strong sense of mastery and very little retention. It is the same mechanism at work in the gap between recognition and recall, where seeing an answer and thinking "yes, I knew that" is a completely different act from producing it from nothing.
The reliable move is to replace the estimate with a measurement. Not a feeling about whether you know it. An attempt to produce it. That is the entire logic of self-testing, and the evidence behind the testing effect is among the sturdiest in learning science, precisely because it does not ask you to introspect. If you want a practical version of the question, knowing when you have studied enough is a decision that gets better the more you replace judgement with retrieval.
Understanding your own thinking as a skill you can develop, rather than a fixed instrument you are stuck with, is the subject of metacognition. The research above should make you neither smug nor despairing about it. People have moderate insight. Moderate insight improves with better information and does not improve much from being told to be humble.
One more thing, and it applies with force to how this term gets used in public. The effect was never about intelligence. It was about calibration on a specific task by people who had just taken a specific test. Using it as a synonym for stupidity gets the science wrong twice over: once because the original studies said nothing about general intelligence, and again because the pattern's clearest surviving form appears among people who know a fair amount, not among people who know nothing. If you want a sense of how genuine expertise develops and what it feels like from the inside, the psychology of expertise is a better place to start than a curve with a mountain on it.
Why This Idea Was So Easy To Believe
It is worth asking why this particular finding escaped the journals and became a household phrase, when hundreds of equally solid results never leave the library.
Part of the answer is that it flatters the person using it. The effect is almost always applied to somebody else. You have probably never heard a person say that they themselves are currently on the peak of Mount Stupid, because the claim is structured so that noticing it is proof you are past it. A theory that cannot be turned on the speaker is a very comfortable theory to own.
Part of it is that it gives a scientific-sounding name to something everyone has felt. We have all met the confident beginner. The label arrived, it fit, and the fit felt like evidence.
And part of it is the picture. A curve with a mountain and a valley tells a story in one glance: a journey, a humbling, a redemption. The real figures from the paper tell no story at all. They show two lines going up. Nobody shares that.
The lesson is not that people are gullible. It is that a memorable image plus a plausible name will outrun a careful result every time, and that the outrunning happens fastest when the result is about other people's flaws.
What Would Settle It
One test of whether an argument is scientific rather than tribal is whether each side can say what would change its mind. Both sides here can, which is a good sign for a dispute that has run for twenty-seven years.
The artifact position would take real damage from an effect that survives methods immune to regression. If self-assessment error reliably widened at low ability under tests built to detect exactly that, across large samples and several domains, the claim that the pattern is mostly a drawing convention would not hold. Some studies already find fragments of this. What is missing is consistency.
The psychological position would take real damage from the opposite result, and it has already absorbed some. A Registered Report found that the quality of a person's metacognitive processing was unrelated to their skill, which is close to the centre of the dual burden claim. More findings like that, in other domains, would leave the original explanation with very little to do.
The most useful evidence would come from designs that never build the flawed chart in the first place. Measure ability with one instrument and belief with another. Use a standing external number, the way the chess study used ratings, so the sorting variable and the benchmark are not the same measurement taken twice. Follow the same people over time so that a change in belief can be observed rather than inferred from a snapshot.
Those studies are harder and slower than handing a quiz to a class and drawing quartiles, which is precisely why the quiz version dominated for two decades. The convenient method sets the direction of a field more often than anyone admits.
There is one more thing worth wanting, and it is not a study. It is a redrawing. As long as the first image a person meets on this topic is a mountain that appeared in no paper, the public conversation will keep being about a picture rather than about a question. The question is a good one. People are moderately good at knowing what they know. Working out exactly how good, and what improves it, is worth more than any curve with a joke written on it.
What The Argument Actually Is, In 2026
Twenty-seven years on, here is the position an honest reader can defend.
The original studies happened, the numbers in them are real, and nobody is accusing anyone of fraud. Four studies of undergraduates found that bottom-quartile scorers rated themselves far above their measured standing and top-quartile scorers rated themselves below it.
The standard way of displaying that finding cannot distinguish a real effect from an artifact, because it puts the sorting variable on both axes and because the ends of a bounded scale constrain which direction errors can point. Random numbers run through the same conventions produce the same picture.
The dual burden explanation, which is the reason the effect became famous rather than merely interesting, has been tested with methods designed for the job and did not survive. Poor performers in that test were appropriately less confident than good performers. The pattern came from the performance scores.
When better methods are applied, the effect shrinks dramatically and in some datasets vanishes. In others it persists in reduced form. The people who have done most to shrink it describe the remaining effect as small and limited rather than absent.
Something real does survive. Thousands of chess players with two decades of objective public feedback are still overconfident about their own strength by a consistent margin. Beginners with a little knowledge are more overconfident than people with none. Almost everyone rates themselves above average on tasks that feel easy.
And the chart you have seen, the one with the peak and the valley and the mountain named after an insult, was never in the paper at all. Somebody drew it. It spread. It got attached to a real study by repetition until it became the thing most people believe the study showed.
There is a decent joke in there about a picture that misrepresents a finding about people misjudging things, and the joke has been made many times. The more interesting point is quieter. A field spent two decades arguing about human nature, and a large part of what it was arguing about turned out to be a property of a graph. That is not a scandal. That is science working slowly, in public, with real names on both sides of the argument, which is exactly what it is supposed to look like.
The next time someone shows you that curve, you will know three things. It is not from the paper. The paper's own chart cannot settle the question. And the answer, so far as anyone can currently say, is that people are moderately good at knowing what they know, and slightly worse at it than they think.
Frequently Asked Questions
What is the Dunning-Kruger effect in simple terms?
It is the finding that people who score lowest on a test tend to rate their own performance far above where it actually falls, while the highest scorers tend to rate themselves slightly below where they actually fall. It comes from a 1999 study by Justin Kruger and David Dunning that ran four experiments on 334 Cornell undergraduates using humour, logical reasoning and grammar. In the humour study, the bottom quarter of scorers were at about the twelfth percentile and estimated themselves at the fifty-eighth. The popular version of the effect, that confident people are incompetent and competent people are humble, is a distortion. The original data showed both groups guessing closer to average than they really were.
Has the Dunning-Kruger effect been debunked?
Not exactly, and the honest answer is more interesting than either headline. What has been seriously damaged is the standard method used to show it. Sorting people into quartiles by test score and then plotting their self-ratings against that same score puts one variable on both axes, and researchers have reproduced the classic pattern using randomly generated numbers with no people involved. A 2022 Registered Report using modern measures of metacognition found the pattern was driven by performance scores rather than by any deficit in self-knowledge, and that poor performers were appropriately less confident than good performers. But a 2025 study of 3,388 tournament chess players found real, persistent overconfidence despite years of objective public feedback. The instrument is broken. Something smaller than the original claim is still there.
Why does the Dunning-Kruger effect appear in random data?
Two reasons stack on top of each other. The first is that the test score is used twice: once to sort people into groups and once as the benchmark those groups are compared against. Correlating a variable with itself produces a relationship whether or not anything real is happening. The second is that percentile scales have ends. Someone at the third percentile has almost no room to underestimate themselves, so their errors can only point upward, and someone at the ninety-seventh percentile has the opposite constraint. Add ordinary measurement noise and regression to the mean, and the two converging lines appear without any human psychology involved.
Is the Mount Stupid curve the real Dunning-Kruger graph?
No. The 1999 paper contains four figures and all of them plot perceived ability and perceived test performance against actual test performance, quartile by quartile. Both lines on those figures rise from left to right. There is no peak, no valley, and no confidence axis. The curve with Mount Stupid and the Valley of Despair was drawn later by other people and became attached to the study through repetition. It is probably the most widely shared image in popular psychology that does not appear in the work it is credited to.
If I cannot trust my own sense of how much I know, what should I do instead?
Replace the estimate with a measurement. A feeling of familiarity is a poor guide, because recognising material and being able to produce it are different abilities, and re-reading something until it feels smooth reliably inflates confidence without adding much retention. Testing yourself gives you an actual data point instead of an impression. It is also worth knowing that feedback alone does not fix miscalibration: students given repeated practice tests with feedback stayed miscalibrated, and chess players with decades of precise public ratings still overestimate themselves. The useful lesson is not to distrust yourself globally. It is to stop asking how confident you feel and start asking what you can currently produce.




