Introduction

Show a chess grandmaster a board for five seconds and take it away. She will rebuild it almost perfectly, twenty-five pieces in their exact squares. Now scramble those same pieces into positions that could never occur in a real game, and show her again for five seconds. Her memory collapses to roughly the level of a beginner [1]. Nothing about her brain changed in those few seconds. What changed was whether the board contained meaning.

That single experiment sits at the centre of the psychology of expertise, and it points to something counterintuitive. Experts are not people with better memories. They are people whose memories have been reorganised by years of exposure to one particular slice of the world. The advantage lives in the material, not in the machinery.

For thirty years, one explanation dominated. Practice enough, in the right way, and anyone can get there. Then the data arrived. A meta-analysis pooling 88 studies found that deliberate practice accounted for 26 percent of the differences in skill among competitive game players, 21 percent among musicians, and 18 percent among athletes [2]. Substantial. Also nowhere near everything.

This article follows that argument from a Dutch psychologist watching chess players think in 1946, through a laboratory where a student trained himself to hold eighty numbers in mind, into brain scanners that caught London taxi drivers growing new grey matter, and out the other side into the uncomfortable question nobody likes asking. What happens when an expert is confidently, catastrophically wrong?

Abstract chessboard with glowing geometric clusters in indigo and amber.

The Chess Master Who Searched No Deeper Than Anyone Else

Adriaan de Groot expected to find calculation. In the 1940s, this Dutch psychologist and competitive chess player set out to discover what separated grandmasters from strong club players, and the obvious hypothesis was depth. Masters must see further ahead. More moves, more branches, more variations.

His method was unglamorous and it worked. He asked players at every level, from world championship contenders down to strong amateurs, to think aloud while solving the same set of positions. Every spoken word was transcribed. Then he counted the moves each player considered, how far ahead each line was calculated, and how much total time went into each phase of the search.

The result surprised him. Grandmasters did not search dramatically deeper than expert players. They examined roughly the same number of possible moves. What they did differently happened in the first few seconds, before any calculation began. They saw the right move almost immediately, then spent their time verifying it. Weaker players generated worse candidate moves and then analysed them carefully. The bottleneck was not thinking. It was seeing.

This flipped the assumption that had guided training for a century. If the difference were search depth, the obvious remedy would be to calculate more. If the difference is in which moves even occur to you, calculation cannot fix it. You cannot analyse a move you never considered.

De Groot also ran a memory test almost as an afterthought. Show a position for a few seconds, then ask for a reconstruction. Masters were near perfect. Ordinary players were not. That side experiment turned out to matter more than the main study.

Twenty-seven years later, two researchers at Carnegie Mellon picked it up and built a theory on it. William Chase and Herbert Simon tested a master, a class A player, and a beginner on both real game positions and randomly arranged ones [1]. On real positions, the master dominated. On random positions, the gap shrank until it nearly vanished.

Their explanation introduced a word that has stuck for fifty years. Chunking. A chunk is a group of items that memory treats as one unit, the way a phone number becomes three blocks rather than ten separate digits. Chase and Simon argued that masters do not store twenty-five pieces. They store five or six familiar configurations, each containing several pieces, and each already known from thousands of previous games.

They even found a way to measure chunk boundaries. When players reconstructed a board, the pauses between placing pieces were not random. Pieces belonging to the same chunk were placed rapidly, in bursts. Then came a longer gap before the next burst began. Chase and Simon set the boundary at roughly two seconds of hesitation, and that two second inter-chunk interval became a standard tool in expertise research.

Here is where the story gets more interesting than the textbook version admits. The strong claim, that expert memory advantage disappears entirely on random material, turned out to be wrong. Fernand Gobet and Simon himself went back and checked it in 1996 with much larger samples [3]. Masters kept a small but reliable edge even on positions that could never occur in a real game. Their average chunk on random material contained about 3.6 pieces, against roughly 2.7 for weaker club players. A companion study confirmed the same skill effect using rapidly presented random positions [4].

Why does a residual advantage matter? Because it means expertise is not purely a lookup table of memorised patterns. Something more flexible is happening, something that can find partial structure even in chaos. A random board is not truly random to a master, because scattered pieces still form fragments of relationships she has seen ten thousand times. A knight and a pawn happen to sit in a configuration that occurs in real games, and that fragment gets encoded as a unit even though the rest of the board is nonsense.

That gap in the theory drove the next twenty years of research.

There is a practical lesson buried in the random board result, and it applies far outside chess. Expert perception is not a general talent that can be switched on in a new field. It is a property of the material. A radiologist who reads chest films with uncanny speed has no advantage whatsoever reading a balance sheet. Meaning has to be built domain by domain, from the ground up, and there is no shortcut through it.

Abstract shapes in glowing indigo clusters versus dim, unconnected forms.

Fifty Thousand Chunks and a Number Nobody Counted

If experts store chunks, an obvious question follows. How many?

Simon and Kenneth Gilmartin tried to answer it in 1973 by building a computer program called MAPP that simulated chess memory [5]. They fed it chunks, tested its recall, and extrapolated. Their estimate for master level performance came out as a range. Somewhere between 10,000 and 100,000 chunks.

That range then did something ranges often do in popular science. It collapsed into a single number.

The figure of 50,000 chunks now appears everywhere, quoted as though somebody counted them. Nobody did. It is a midpoint approximation drawn from a simulation, offered by its own authors as an order of magnitude rather than a measurement. Gobet and Simon later phrased it carefully, saying that recall close to master level requires at least 50,000 chunks, which is a floor rather than a fixed inventory. Anyone repeating the number as hard fact has lost the original uncertainty along the way.

Correcting that matters for a practical reason. A precise number invites a precise plan, as though skill were a warehouse with a known capacity. A range spanning one order of magnitude tells a different story. It says the architecture varies, the counting method is contested, and nobody has opened a brain and tallied its contents.

Meanwhile, chunk theory had a problem it could not solve. Chunks are supposed to live in short term memory, which holds only a handful of items. Yet strong players can reconstruct several boards presented one after another, far beyond any plausible short term capacity.

Gobet and Simon answered this in 1996 with template theory [6]. A template is a chunk that has been used so often it has hardened into a structure with slots. Think of it as a form with blanks. The general shape is already stored, and the specific details of today's position simply drop into the empty fields. Filling a slot is fast. Building a new structure is slow. That difference explains how a master can encode a whole board in seconds while a novice is still counting pieces.

They tested it directly. In their 1998 study revisiting the chunking hypothesis, a chess master was shown multiple boards in sequence [7]. He reconstructed up to nine positions with high accuracy, which is roughly 160 pieces. No version of short term memory can hold that. Templates, sitting in long term memory and accessed through rapid indexing, can.

A later review compared four competing theories of expert memory and concluded that template theory handled the widest range of findings, though not without unresolved edges [8]. Recent work using chess tasks of graded complexity has continued to refine how perceptual expertise scales with difficulty [9].

The template idea also explains something coaches have noticed for generations without being able to name it. Beginners in any field ask for rules. Intermediates ask for more rules. Advanced practitioners stop asking, because the question no longer makes sense to them. A template is not a rule. It is a recognised situation with variable parts, and no verbal instruction can install one. Only exposure can.

What does this mean outside chess? It means that when a skill starts feeling easy, the change is usually not that you are thinking faster. It is that you have stopped assembling things from parts. The parts arrive pre-assembled.

It also predicts where the fluency will break. Templates are built from the situations you have actually encountered. Present an expert with a case that superficially resembles a familiar template but differs in a critical slot, and the template will fire anyway. This is not a rare edge case. It is the standard failure mode of experienced practitioners, and later sections return to it.

No

Yes

Raw Perception

Familiar Pattern?

Piece by Piece Encoding

Chunk Retrieved

Template With Slots

Long Term Storage

Working Memory Overload

The diagram above compresses decades of argument into one pathway, and the branch on the left is where most learners spend their time. Encoding piece by piece is not a character flaw. It is what every brain does before the patterns exist.

Transparent geometric spheres with intricate lattice structures in amber and blue.

The Student Who Trained His Memory to Hold Eighty Numbers

In 1978, an undergraduate at Carnegie Mellon agreed to a tedious experiment. Sit in a lab, listen to random digits read aloud at one per second, and repeat them back. His starting span was seven digits, which is completely ordinary.

Anders Ericsson and William Chase kept him coming back. One hour a day, three to five days a week, for more than two years. Over 230 hours of practice.

By the end, the student identified in the literature as SF could repeat back 79 digits [10]. Later sessions pushed him to 82. A second participant, trained for around 800 hours, exceeded 100.

The obvious interpretation is that practice expanded his short term memory. It did not. That is the finding that makes the study famous.

SF was a competitive distance runner. Somewhere in the first weeks he started hearing digit groups as running times. The sequence 3492 stopped being four digits and became three minutes and forty-nine point two seconds, a near world record mile. Groups of four became times, times were grouped into larger sets, and the whole structure was indexed by categories he already knew intimately. His short term memory still held three or four groups. The groups had simply become enormous.

The proof came from a control test. When the researchers switched from digits to consonants, SF's span dropped back to about six. He had not built a better memory. He had built a better filing system for one specific kind of material, and it did not transfer.

This connects directly to a limit that had been assumed for decades. George Miller's 1956 paper put the capacity of immediate memory at seven items, plus or minus two [11]. Nelson Cowan revised that downward in 2001, arguing that when rehearsal and grouping strategies are properly controlled, the real limit is closer to four chunks [12]. Four. Not seven.

Expertise does not raise that ceiling. It changes what fits underneath it. Anyone who has studied how working memory limits shape learning will recognise the pattern, because the same constraint drives instructional design.

Ericsson and Walter Kintsch went further in 1995 with the theory of long term working memory [13]. Their argument was that skilled performers build retrieval structures, essentially cue systems that allow information stored in long term memory to be accessed almost as fast as if it were held in mind. A waiter who takes twenty orders without writing anything down is not straining short term memory. He is dropping each order into a spatial map of the room.

The same principle applies to reading. A skilled reader of a technical paper is not holding every sentence in mind. She is dropping each new claim into an existing structure of what the field argues about, which is why she can be interrupted mid paragraph and resume without rereading. A novice loses the thread because there is no structure to drop things into, so every sentence has to be held actively until it can be connected to something.

The theory has critics. Gobet argued in 2000 that the concept of a retrieval structure quietly merges three different kinds of memory organisation that behave differently, and that template theory explains the same data with fewer assumptions [14]. The disagreement is still live.

The disagreement is also narrower than it looks from outside. Nobody in this debate thinks experts have bigger short term memories. The argument is about the shape of the machinery that lets long term storage behave, temporarily, like working memory. Whether that machinery is best described as a retrieval structure or a filled template is a real scientific question with real consequences for how skill acquisition is modelled. It is not a question about whether the phenomenon exists.

TheoryUnit of storageWhere it livesStrongest evidence
Chunk theory (1973)Small perceptual groupShort term memory, indexed to long termRandom board collapse and two second pause boundaries
Template theory (1996)Structure with variable slotsLong term memory, rapidly filledRecall of up to nine chess boards in sequence
Long term working memory (1995)Domain specific retrieval structureLong term memory with fast cue accessDigit span of 79 and skilled reading comprehension

Notice that all three theories agree on one thing while disagreeing on mechanism. None of them says the expert has more raw capacity. Every one of them says the expert has better organisation.

What does this mean for anyone studying something difficult? It reframes what feeling overwhelmed actually indicates. The sensation of too much information is rarely a sign of a weak memory. It is a sign that the material has not yet been organised into units, so every element is competing for the same four slots. The fix is not more effort at holding things in mind. The fix is finding or building the structure that lets several elements travel together.

SF's running times were not a memory trick in the usual sense. They were a bridge from unfamiliar material to knowledge he already possessed in depth. That is the general recipe, and it is why the same new concept lands easily for one person and bounces off another. The difference is what each of them already has to attach it to.

Sorting Physics Problems Reveals How Experts Actually See

In 1981, Michelene Chi, Paul Feltovich and Robert Glaser ran an experiment so simple it is almost annoying. They gave people twenty-four physics problems on index cards and asked them to sort the cards into piles based on similarity of solution [15]. Eight advanced doctoral students in physics. Eight undergraduates who had completed one semester.

Nobody solved anything. They just sorted.

The undergraduates grouped by what the problems looked like. Inclined plane problems together. Pulley problems together. Spring problems together. Reasonable, and completely useless for solving them.

The physicists grouped by the principle needed. Conservation of energy in one pile, Newton's second law in another. Two problems that looked nothing alike sat together because the same law cracked both.

That is the clearest demonstration in cognitive psychology that expertise changes perception rather than just adding facts. The novices were not lazy. They sorted by the only feature they could see. The deep structure was invisible to them because they had no representation of it yet.

The design deserves credit for what it avoided. Because nobody had to solve anything, the study cleanly separated perception from problem solving ability. Any difference in sorting could not be explained by the experts simply being faster or better at calculation. They were categorising the world differently before any calculation began.

The finding has been reproduced across mathematics, programming, biology and law, always with the same shape. Beginners cluster by surface. Experts cluster by function. And in every case, the surface features that mislead novices are exactly the features that textbooks use to organise chapters, which may explain part of why transfer from coursework to real problems is so poor.

Medicine produced a stranger version of the same finding. Henk Schmidt, Geoffrey Norman and Henny Boshuizen mapped how clinical reasoning develops and found something that looked like a mistake in the data [16]. As students advanced, their explicit use of underlying biomedical science went down, not up. Experienced physicians talked less about pathophysiology than intermediate students did.

Boshuizen and Schmidt called the process knowledge encapsulation [17]. The biomedical knowledge does not vanish. It gets packed inside higher level clinical concepts, so that a phrase like unstable angina now silently contains the entire mechanism. The expert can unpack it when a case turns strange. Most of the time she does not need to.

There is a measurable side effect. Ask people to recall the details of a clinical case and performance does not rise steadily with training. It peaks in the middle. Intermediate students remember the most detail, because they are still processing everything explicitly, while experts have already compressed the case into a pattern and discarded the surface [18]. This intermediate effect is one of the reasons students often outperform their supervisors on detail recall tests and then lose badly on diagnosis.

The way clinical knowledge gets restructured during training follows exactly this trajectory, and it explains a frustration many learners report. The moment you start understanding a field is often the moment you stop being able to recite it.

What does this mean for anyone learning something hard? Losing the ability to list every detail is not decay. It can be a sign that the details have been absorbed into something larger. The test is whether you can still unpack them on demand.

The Number That Escaped the Laboratory

Now the most misunderstood study in the psychology of expertise.

In 1993, Anders Ericsson, Ralf Krampe and Clemens Tesch-Römer published a paper in Psychological Review that would be cited more than nine thousand times [19]. They studied violinists at the Music Academy of West Berlin, sorted by their professors into three groups. The best students, headed for solo careers. Good students. And students on the music teacher track, the least accomplished of the three.

The method matters here. The researchers used retrospective estimates and detailed diaries, asking players to reconstruct their weekly practice going back to the age they first picked up the instrument. Accumulated hours were then calculated for each year of life.

The pattern was clean. By age twenty, the best group had accumulated roughly 10,000 hours of solitary practice. The good group had roughly 2,500 hours less. The teacher track group was around 5,000 hours behind the best.

Ericsson's argument was not that hours matter. It was that a specific kind of practice matters, which he called deliberate practice.

The definition was strict, and most people who quote the theory have never read it. Deliberate practice requires a task chosen specifically to improve a weakness that has already been identified. It requires operating just past the edge of current ability, where errors are frequent. It requires immediate and informative feedback, either from a teacher or from the task itself. It requires repetition with refinement rather than repetition alone. And it is effortful and generally not enjoyable, which is why almost nobody sustains more than a few hours of it per day.

Playing a piece you already know is not deliberate practice. Playing casual games is not deliberate practice. Reading about a skill is not deliberate practice. By Ericsson's definition, most of what people call practice does not qualify.

He reviewed the broader evidence with Andreas Lehmann in 1996 and extended the framework to medicine and other professions in later work [20], [21], [22]. The medical papers argued something genuinely useful, which is that years of clinical experience are not equivalent to years of structured skill development, and that most professional work provides almost none of the second.

Then in 2008 a journalist read the paper.

Malcolm Gladwell's Outliers turned that 10,000 hour average into a rule, and the rule into a promise. Ten thousand hours was presented as the magic number of greatness, the threshold that separates the extraordinary from everyone else.

This is the first myth worth killing properly, because almost every article on this topic still repeats it. The ten thousand hour rule is Gladwell's, not Ericsson's. Ericsson spent the rest of his career objecting to it.

His objections were specific. The figure was a group average, and roughly half the violinists in the top group had not reached 10,000 hours by age twenty. There is no threshold in the data, no point where something switches on. The number varies enormously across domains, and in some fields elite performance arrives far sooner. Most importantly, Gladwell's version dropped the word deliberate, which is where the entire theory lived. Ten thousand hours of comfortable repetition produces ten thousand hours of comfortable repetition.

Ericsson made the point sharply in his 2016 reply to critics, arguing that the field had begun measuring the wrong thing entirely by counting general practice rather than the specific structured activity he had described [23].

He was right that the number had been mangled. He was about to discover that the underlying claim had its own problems.

1894
Binet studies memory in blindfold chess players
1946
De Groot finds masters search no deeper than others
1973
Chase and Simon publish the chunking theory of skill
1980
A student trains his digit span from seven to seventy-nine
1993
Ericsson links elite performance to deliberate practice
1996
Gobet and Simon replace rigid chunks with flexible templates
2008
Outliers popularises the ten thousand hour rule
2011
Taxi driver study shows training reshapes the hippocampus
2014
Meta-analysis finds practice explains a minority of skill
2026
Elite chess data shows deliberation still beats pure intuition

Reading that sequence in order reveals something the popular version hides. The practice theory arrived in the middle of the story, not at the end, and the twenty years after it were spent testing it rather than celebrating it.

Worn stone staircase with smooth hollows, illuminated by dusty sunlight.

What Twenty Years of Data Did to a Beautiful Idea

Brooke Macnamara, David Hambrick and Frederick Oswald did the arithmetic in 2014 [2]. They gathered every study they could find that measured both accumulated practice and performance, pooled 88 of them, and calculated how much of the variation in skill practice actually explained.

The answers, by domain, were 26 percent for games, 21 percent for music, 18 percent for sports, 4 percent for education, and less than 1 percent for professions.

Read those numbers slowly. In chess and similar games, the domain where the practice theory looks strongest, roughly three quarters of the differences between people were something other than how much they had practised. In professional work, practice explained almost nothing measurable.

Share of skill differences explained by accumulated practiceGamesMusicSportsEducationProfessions302826242220181614121086420Percent of variance

One note on that chart. The professions bar is drawn at one percent so it remains visible, but the measured value fell below one. The honest reading is that in the messiest domains, accumulated hours told researchers almost nothing about who was good.

A parallel re-analysis by Hambrick and colleagues in the journal Intelligence reached the same conclusion from a different direction, finding that in chess and music deliberate practice accounted for around a third of the reliable variance at most [24]. A later meta-analysis focused specifically on sport reported a similar picture [25].

Then came the study that unsettled people most.

Miriam Mosing and colleagues in Sweden had access to something no laboratory could build. A twin registry. They surveyed 10,500 Swedish twins on music practice and tested music ability [26]. Practice and ability correlated, as expected. But two other findings landed harder.

First, the propensity to practise was itself substantially heritable, with estimates between 40 and 70 percent. Whatever makes a person willing to grind through unpleasant repetition is not purely a choice made in a vacuum.

Second, and this is the uncomfortable part, when they compared identical twins who differed in how much they had practised, the twin who practised more was not reliably better. In one pair the difference was 20,228 hours. Twenty thousand extra hours, no detectable advantage in measured ability.

Design matters here, so it is worth being precise. A twin comparison controls for genetics and shared upbringing, which is exactly what a correlational study cannot do. It does not prove practice is useless. It does show that the simple causal story, more practice therefore more ability, does not survive the controls.

In 2019 Macnamara and Megha Maitra went back to the original 1993 study and tried to reproduce it [27]. The relationship between practice and skill group was still there and still substantial, but weaker than in the original. More awkwardly, the specific claim that teacher designed solitary practice was the critical ingredient did not hold up as cleanly as the theory required.

So what replaced it? Not a swing back to innate talent. Something more careful.

Elizabeth Meinz and Hambrick had already shown that working memory capacity predicted sight reading ability in pianists even after accounting for practice hours [28]. Fredrik Ullén, Hambrick and Mosing proposed a multifactorial gene-environment interaction model, in which genetic and environmental influences on ability, on personality, and on the willingness to practise all feed into each other over years [29]. Hambrick and colleagues summarised the shift as moving beyond born versus made altogether [30], and the edited volume on the science of expertise gathered the behavioural, neural and genetic strands into one framework [31].

What else predicts skill, then? Several things, none of them decisive on its own.

General cognitive ability shows up most strongly in the early phases of learning, when a task still has to be worked out consciously, and its predictive power tends to shrink as performance automates. Working memory capacity keeps predicting in domains where the task never fully automates, such as sight reading unfamiliar music. Age of starting matters, though disentangling a biological sensitive period from the simple fact that early starters accumulate more years remains difficult. Personality traits associated with sustained effort predict how much practice happens in the first place, which is exactly the pathway the twin data pointed at.

And physical factors matter in the domains where they matter, which sounds obvious until someone tries to explain elite basketball through practice hours alone.

The most recent statement of where the field stands came in December 2025, when Kathryn Friedlander summarised the argument in the Journal of Expertise [32]. Her framing is that expertise research spent decades trapped in a narrow fight between deliberate practice and innate aptitude, and that multifactorial models have finally opened it up.

StudyYearDesignHeadline result
Ericsson Krampe and Tesch-Römer1993Retrospective diaries with 3 skill groupsBest violinists near 10000 hours by age 20
Macnamara Hambrick and Oswald2014Meta-analysis of 88 studiesPractice explains 26 percent in games and under 1 percent in professions
Mosing and colleagues2014Twin study with 10500 participantsPractice itself 40 to 70 percent heritable and no causal effect within twin pairs
Macnamara and Maitra2019Replication of the 1993 designEffect present but smaller and the teacher practice claim did not hold

What does this mean for someone trying to get good at something? The practical advice survives almost intact. Structured, effortful, feedback rich practice still beats comfortable repetition by a wide margin, in every domain measured. What does not survive is the guarantee. Practice is the part you control. It is not the only part that matters.

Identical seedlings in pots growing at different rates under lamps.

Brains That Grew Because of What They Learned

London taxi drivers memorise twenty-five thousand streets. The qualification, known simply as the Knowledge, takes three to four years and most candidates fail.

In 2000, Eleanor Maguire and colleagues at University College London scanned sixteen right handed male taxi drivers and compared their brain structure to matched controls [33]. The posterior hippocampus, the rear portion of a seahorse shaped structure buried in each temporal lobe that builds spatial and episodic memory, was significantly larger in the drivers. The anterior portion was smaller. Volume in the posterior region correlated with years spent driving.

A follow up compared taxi drivers to bus drivers, controlling for driving experience, stress and time behind the wheel [34]. The difference held. It was not driving. It was navigating.

But both studies shared a fatal ambiguity. Correlation. Perhaps people with larger posterior hippocampi are drawn to a job that rewards spatial memory, and the brain difference came first.

Katherine Woollett and Maguire settled it in 2011 with a design that took four years to run [35]. They scanned 79 trainee taxi drivers at the start of their training and again three to four years later, alongside 31 controls. At the outset, no structural difference existed between any of the groups.

By the end, 39 trainees had qualified and 40 had not. Grey matter in the posterior hippocampus had increased only in those who qualified. The trainees who dropped out or failed looked like the controls. Same starting point, different outcome, and the difference tracked the learning rather than predicting it.

That is as close to causal evidence as human structural imaging gets. Learning changed the tissue.

There was a cost, and reporting it matters. The qualified drivers, who had gained spatial memory for London, performed worse than controls on a separate task involving learning and recalling complex visual figures. The advantage was not free. Something else appeared to give way as the spatial system expanded, which fits the broader theme of this article rather neatly. Expertise is a reallocation, not an addition.

The same structure supports how the hippocampus selects what gets stored, which is why a job built entirely on spatial recall shows up there rather than somewhere else.

Musicians produced parallel findings. Sara Bengtsson and colleagues used diffusion tensor imaging, a technique that measures how water molecules move along nerve fibres and therefore how well organised the white matter is, to compare professional pianists with non-musicians [36]. Practice hours correlated with fibre organisation, and the pattern differed by life stage. Childhood practice related to the widest set of regions including the pyramidal tract, the main highway carrying movement commands from cortex to spinal cord.

That life stage pattern points at something important. The regions whose organisation correlated with adult practice were fewer and more restricted than those tied to childhood practice. One reading is that certain white matter pathways are still actively maturing during childhood and adolescence, so training during that window shapes them in ways later training cannot. Another reading is simply that early starters practise longer in total. The design cannot separate the two, and the authors did not claim it could.

Jan Scholz and colleagues showed the same thing prospectively by teaching people to juggle and scanning them before and after, finding white matter changes in the intraparietal sulcus [37]. At the cellular level, work in mice demonstrated that generating new myelin, the fatty insulation that speeds signal conduction along axons, is required for normal motor skill learning [38]. Christian Gaser and Gottfried Schlaug found grey matter differences between professional musicians, amateurs and non-musicians that scaled with training status [39].

Now the third myth, and this one is still being argued about.

In 2000, Isabel Gauthier and colleagues scanned bird experts and car experts and found that objects in their domain of expertise activated the fusiform face area, a patch of cortex on the underside of the temporal lobe that responds strongly to faces [40]. The conclusion drawn by many was that this region is not about faces at all. It is a general expertise module.

That conclusion is not settled. Yi Xu revisited the question in 2005 and found the expertise effect was real but considerably smaller than the face response, and inconsistent across measures [41]. Nancy Kanwisher has argued at length that the region shows genuine domain specificity for faces which expertise effects do not explain away [42]. Merim Bilalić and colleagues examined chess experts and found expertise related activity in neighbouring regions rather than a simple takeover of the face area [43].

The accurate statement is that visual expertise recruits parts of the ventral temporal cortex that overlap with face processing, and that whether this reflects one shared mechanism or two adjacent ones remains open. A review of physical and mental training effects on the brain reaches a similarly cautious conclusion across domains [44].

Abstract layered contour relief with indigo and cream paper cut style.

The Honest Limits of What Brain Scans Can Show

A section like the one above is where science writing usually stops. It should not.

Structural brain imaging of expertise has known weaknesses, and pretending otherwise would misrepresent the evidence. Most studies are cross sectional, comparing experts to non-experts at a single moment, which cannot separate the effects of training from the traits that led someone into the training. Sample sizes are often small, frequently under thirty per group, which makes findings fragile.

Cyril Pernet and colleagues, and separately Marcus Boekel and colleagues, have shown how easily structural brain and behaviour correlations fail to replicate. Boekel's team attempted to reproduce seventeen previously published structural associations and found convincing support for almost none of them [45]. Methodological reviews have also flagged how sensitive voxel based morphometry results are to analysis choices [46].

This is why the Woollett and Maguire longitudinal design matters so much. It is one of the few studies in this literature that can support a causal claim, and it is repeatedly cited precisely because so few others can.

What would strengthen the field is not complicated, only expensive. Pre-registered longitudinal designs with sample sizes several times larger than current norms, standardised analysis pipelines agreed in advance, and comparison groups matched on the traits that draw people into a discipline in the first place. Until more of that exists, the correct posture towards any headline claiming that a particular activity grows a particular brain region is interest rather than belief.

None of this means the effects are imaginary. It means the confidence intervals in popular reporting are considerably narrower than the confidence intervals in the data.

The functional side rests on firmer ground. As a motor or cognitive skill becomes automatic, activity reliably shifts away from prefrontal regions involved in effortful control and towards the basal ganglia and cerebellum, structures deeper in the brain that handle sequenced and habitual action. Russell Poldrack and colleagues traced this shift directly as participants learned a skill to automaticity [47]. Julien Doyon and Habib Benali described the two parallel circuits involved in motor sequence and motor adaptation learning [48], and Gregory Ashby and colleagues modelled how the striatum trains cortex during automatisation [49].

That shift is what people are describing when they say a skill has become second nature, and the same transition underlies how repeated behaviours become automatic in everyday life.

What does this mean practically? It means the feeling of effort dropping away is a real biological event, not a mood. It also means that once a skill has moved into those circuits, deliberately re-examining it becomes genuinely harder, which sets up the next problem.

Minimalist cross section of translucent glass plates with layered details.

Good Moves That Block Better Ones

In 1942, Abraham Luchins gave people a series of water jar problems [50]. Given jars of three different capacities, measure out an exact amount. The first several problems all required the same three step formula. Fill the large jar, pour off the medium, pour off the small twice.

Then he slipped in problems that could be solved that way, but could also be solved in one obvious step.

Around 81 percent of participants who had been trained on the long method kept using it. A control group that skipped the training solved the easy problems directly almost without exception. The trained group was not less intelligent. They were more prepared, and being prepared cost them.

Luchins called it Einstellung, the German word for set or attitude. A mental set that has been rewarded stops being a tool and starts being a lens.

Sixty-six years later, Merim Bilalić, Peter McLeod and Fernand Gobet built a chess version and added eye tracking, which turns an argument about self awareness into a measurement [51]. They presented positions containing a familiar winning pattern, the smothered mate, alongside a faster and better solution.

When the familiar solution was available, strong players found the optimal one far less often than they did on control positions where only the optimal solution existed. That is the effect. But the eye tracking is the part that unsettles people.

Players who had spotted the familiar solution reported that they were now searching for something better. Their eye movements said otherwise. Gaze kept returning to the squares and pieces relevant to the first solution. They were sincerely convinced they were exploring. They were not.

The companion paper quantified how the effect scales across skill levels, showing that Einstellung shifts performance downward by roughly the equivalent of a full skill class [52]. A grandmaster under Einstellung conditions performs like an international master. An international master performs like a candidate master. Expertise does not protect against this. It supplies the ammunition.

Later work by Heather Sheridan and Eyal Reingold refined the boundary conditions [53]. Using a version where the familiar move was an outright blunder, they found that experts disengaged from it quickly and reliably. The trap only works when the first solution is genuinely good. Merely adequate is the dangerous zone.

That boundary condition has an unpleasant implication for organisations. Obviously bad ideas get caught. Good ideas that happen to be second best sail straight through, and they sail through faster in teams where everyone shares the same training and therefore the same pattern library. A room full of people who all trained the same way is a room where the same familiar solution fires simultaneously in every head, and the resulting agreement feels like validation.

Luchins noticed a version of this in 1942. When he added a warning to the instructions, telling participants explicitly not to be blind, the rate of Einstellung responses dropped substantially. A direct warning helped. Simply being smart did not.

Yes

No

Problem Appears

Familiar Pattern Fires

Is It Clearly Bad?

Attention Disengages

Attention Stays Anchored

Better Solution Found

Search Feels Complete

Better Solution Missed

The branch on the right of that diagram is the entire problem in one line. Search feels complete. Not is complete. Feels.

What does this mean for real decisions? It suggests that the most dangerous moment in any expert judgment is not confusion. It is the arrival of a solution that works. That is precisely when a deliberate search for alternatives has to be forced, because it will not happen on its own.

Why Experts Cannot Explain What They Know

There is a cost to compression that shows up the moment an expert tries to teach.

Colin Camerer, George Loewenstein and Martin Weber demonstrated in 1989 that better informed people cannot successfully ignore what they know when predicting the judgments of less informed people [54]. They named it the curse of knowledge. Raymond Nickerson later built a formal model of how people impute their own knowledge to others and showed how systematically the error runs [55].

Pamela Hinds put a number on the practical damage [56]. Experts asked to estimate how long a novice would need to complete a task underestimated badly, and standard debiasing instructions barely helped. Telling someone to account for their expertise does not make them able to.

In education the same effect has a name. Mitchell Nathan and Anthony Petrosino documented the expert blind spot among trainee teachers, showing that those with deeper subject knowledge tended to organise instruction around the formal structure of the discipline rather than around what learners could actually process [57]. They designed lessons for the person they had become, not the person in front of them.

The National Research Council's synthesis of learning research made this one of its central points about expertise, noting that fluent retrieval and highly organised knowledge do not translate into an ability to teach [58]. Those are separate skills that happen to look related.

Think about what the earlier sections predict here. If knowledge has been encapsulated into templates and clinical scripts and chunks, then the intermediate steps genuinely are not available for inspection. The expert is not withholding them. They have been compiled away.

Michael Polanyi captured the same idea decades earlier with a phrase that has outlived most of his philosophy. People know more than they can tell. A skilled practitioner can demonstrate a judgment reliably and still produce, when asked to justify it, a reconstruction that bears little relationship to what actually happened in her head.

That reconstruction problem matters enormously for how professions train people. If expert self reports are unreliable descriptions of expert cognition, then curricula built by asking experts what they do will encode the reconstruction rather than the skill. This is one reason cognitive task analysis, which infers expert reasoning from behaviour under carefully designed conditions, tends to outperform interviews.

This is also why the best explanations often come from people who learned the material recently and painfully. They still have the assembly instructions.

What does this mean if you are trying to learn from an expert? Ask about failures rather than methods. Ask what confused them at your stage. The compiled version of their knowledge is the least transferable thing they own.

Ceramic vessel on a table beside an empty potter's wheel.

When Expert Intuition Deserves Trust and When It Does Not

Two psychologists spent years disagreeing about expertise in public. Daniel Kahneman had built a career documenting the systematic errors in human judgment. Gary Klein had built one documenting how fire commanders and nurses make excellent decisions under pressure using pattern recognition they cannot articulate.

In 2009 they published their answer together, in a paper openly subtitled a failure to disagree [59].

Their conclusion is the single most useful idea in this entire field. Expert intuition is trustworthy when two conditions are met, and unreliable when either is missing.

The first condition is that the environment must be sufficiently regular to be predictable. There have to be stable relationships between cues and outcomes. Chess has them. Firefighting has them. Long range political forecasting does not.

The second is that the person must have had adequate opportunity to learn those regularities through practice with feedback that arrives quickly and unambiguously. A physician treating a condition with immediate visible response learns. A therapist whose patients improve or decline over years, with no clean signal, may not.

Robin Hogarth built the same distinction into a framework of kind and wicked learning environments [60]. In a kind environment, feedback is accurate, immediate and representative of the situations you will face later. In a wicked one, feedback is delayed, missing, biased or actively misleading, and years of experience can therefore teach the wrong lesson with total confidence.

FeatureKind environmentWicked environment
Feedback timingImmediateDelayed or absent
Feedback accuracyClear and unambiguousNoisy biased or misleading
Rule stabilityRegular and repeatingShifting or one off
Effect of experienceSkill reliably improvesConfidence grows faster than accuracy
Typical exampleChess anaesthesiology weather forecastingLong range forecasting stock picking some clinical prediction

The empirical support for the pessimistic half is uncomfortable. William Grove and colleagues meta-analysed 136 studies comparing clinical judgment against simple mechanical prediction formulas [61]. Mechanical prediction was about 10 percent more accurate on average. It was clearly superior in 33 to 47 percent of studies, while clinical judgment was clearly superior in only 6 to 16 percent.

Philip Tetlock ran an even longer test. Over roughly twenty years he collected around 28,000 forecasts from 284 people who made their living commenting on political and economic trends [62]. Average accuracy was famously described as barely better than a dart throwing chimpanzee. The only stable predictor of accuracy was cognitive style. Forecasters who held many small ideas and updated frequently outperformed those committed to a single large framework, and the second group did worse than chance.

The follow up work was more encouraging. Barbara Mellers and colleagues showed in the Good Judgment Project that training in probabilistic reasoning, team collaboration and structured aggregation produced measurably better forecasters [63]. Calibration is trainable. It just does not arrive automatically with experience.

Medicine provides the sharpest example. A systematic review by Niteesh Choudhry, Robert Fletcher and Stephen Soumerai examined the relationship between years in practice and quality of care [64]. In most of the studies reviewed, more experience was associated with the same or lower quality of care, not better. The likely mechanism is not decay of intelligence. It is a wicked feedback environment plus knowledge that ages.

This gap between felt confidence and measured accuracy is closely related to the illusion of knowing, and both come from the same source. Fluency is a terrible signal for correctness.

So how should any of this be applied? The framework converts into a short set of questions that can be asked about any domain, including one's own.

Does the same situation recur often enough to have a stable structure? Does the outcome become visible, and how long does it take? Is the feedback contaminated by the decision itself, as when a rejected job candidate never gets to prove the rejection wrong? And has this particular person actually had thousands of repetitions with visible outcomes, or merely thousands of days on the job?

Where the answers are favourable, expert intuition is often better than any formula anyone has built. Where they are not, the appropriate response is not to distrust experts and trust yourself instead. It is to reach for structure. Checklists, base rates, explicit criteria set before the decision, and independent judgments aggregated afterwards all outperform unaided intuition in exactly the environments where unaided intuition fails.

Two Kinds of Expert and Only One of Them Adapts

In 1986, two Japanese researchers made a distinction that most of the Western literature ignored for twenty years.

Giyoo Hatano and Kayoko Inagaki studied abacus masters, people who could perform arithmetic at extraordinary speed by manipulating a mental image of the beads [65]. Fast, accurate, and by any reasonable measure expert. But when asked to explain why the procedure worked, or to adapt it to a novel problem, many struggled.

Hatano and Inagaki called this routine expertise, and contrasted it with adaptive expertise, where deep conceptual understanding allows the person to invent new procedures when the situation demands it. Both are real expertise. They develop under different conditions.

Their proposed conditions are worth stating. Adaptive expertise is more likely when the work naturally varies so that fixed procedures break down, when the environment is psychologically safe enough that experimentation is not punished, and when the culture values doing the job well over doing it quickly.

That last condition explains a great deal about why organisations that optimise relentlessly for speed tend to produce technicians rather than innovators.

DimensionRoutine expertiseAdaptive expertise
Core strengthSpeed and accuracy on familiar tasksFlexible response to novel problems
Knowledge typeProcedural and highly automatedProcedural plus deep conceptual understanding
Response to noveltyApplies the closest known procedureReconstructs an approach from principles
VulnerabilityBreaks down when conditions shiftSlower on pure routine throughput
Growth conditionRepetition in a stable environmentVariation feedback and psychological safety

Carl Bereiter and Marlene Scardamalia described a closely related mechanism in 1993 and called it progressive problem solving. Their observation was that as a task becomes easier, practitioners face a choice that most of them never notice making. The mental resources freed by automation can be pocketed, leaving the job easier at the same level of quality, or they can be reinvested by redefining the problem at a harder level. Experts who keep improving are the ones who reinvest. Experts who plateau are the ones who pocket.

That framing explains why a decade of experience produces spectacular results in some people and almost none in others working the same job. It is not the hours. It is what happens to the surplus.

Medical education has taken this seriously. Maria Mylopoulos and Nicole Woods argued that training which produces only efficient pattern matching leaves physicians badly equipped for the cases that do not match anything [66]. The goal is not to choose between efficiency and innovation but to develop both.

Stage models tried to capture this progression, and their history is instructive. Patricia Benner adapted the Dreyfus brothers' five stage model to nursing in the early 1980s, describing a path from novice following explicit rules to expert operating on intuitive grasp of whole situations [67]. It became enormously influential in professional training.

It has also been sharply criticised. Gobet and Philippe Chassy argued that the model lacks empirical support for discrete stages, that its account of intuition is psychologically implausible, and that the evidence instead favours gradual accumulation of stored patterns [68]. Gary Klein, who spent his career studying expert intuition and who was involved in the original Air Force research that produced the model, publicly called for retiring it, noting among other things that in many domains beginners have no explicit rules to follow in the first place [69].

Even the shape of the learning curve turned out to be wrong. For decades the power law of practice was treated as settled, describing how performance improves with repetition. Andrew Heathcote, Scott Brown and Douglas Mewhort tested it properly by fitting functions to individual learners rather than averaged data, using 40 datasets covering 7,910 learning series from 475 participants across 24 experiments [70]. An exponential function fit better in essentially every individual case. The power law was an artefact of averaging across people who learn at different rates.

Klein's own contribution, the recognition primed decision model, describes what experts actually do under time pressure. They do not compare options. They recognise the situation as a familiar type, generate one plausible course of action, mentally simulate it, and either run it or modify it [71].

Which brings the story to the most recent data. In 2026, Michael Vidulich and Pamela Tsang analysed archival tournament results from 20 grandmasters and 7 world chess champions playing under different time controls, spanning move bands of 2,934, 2,905, 2,859, 2,849 and 1,250 moves [72]. If elite expertise were purely intuitive pattern recognition, extra thinking time should add little. It did not work out that way. Even the strongest players in history performed better with more time to deliberate.

The comfortable story, that mastery means replacing slow thought with fast recognition, is too simple. Recognition gets the expert to a good starting point. Deliberation is still what separates good from best.

Conclusion

The psychology of expertise has spent eighty years circling one question, and the answer keeps changing shape.

De Groot went looking for deeper calculation and found faster perception. Chase and Simon explained that perception with chunks, and then their own follow up work showed chunks were not enough. Ericsson explained skill through deliberate practice, and then two decades of meta-analyses and twin studies showed that practice, while genuinely important, accounts for a minority of the differences between people. Every clean story in this field has been complicated by the next set of data, which is what a healthy science looks like.

What survives is a picture with three layers. Expertise is built out of memory structures, chunks and templates and retrieval systems, that let a practised brain see meaning where an unpractised one sees noise. Those structures are physically real, visible as changes in grey matter and white matter organisation, though the imaging evidence is thinner and more fragile than popular accounts suggest. And they carry costs that scale with their benefits. The same compression that makes an expert fast makes her a poor explainer. The same pattern library that supplies the good move blocks the better one.

The most useful thing the field has produced is not a training method. It is a diagnostic question. Does this domain give clear, quick, honest feedback? In domains that do, trust the expert's intuition, including your own. In domains that do not, experience will still generate confidence. It just will not generate accuracy.

There is a second question worth carrying alongside it. When a task gets easier, where does the freed capacity go? Into finishing sooner, or into attempting something harder? That choice, made quietly and repeatedly over years, appears to separate the practitioners who keep growing from the ones who stopped growing a decade ago without noticing.

Expertise, in the end, is not a quantity of knowledge. It is a shape that knowledge takes. And shapes, however elegant, always fit some problems better than others.

Frequently Asked Questions

Is the 10000 hour rule true?

No. The figure came from a 1993 study where the best violinists averaged around 10000 hours by age twenty, but it was a group average with wide variation and no threshold effect. Malcolm Gladwell turned it into a rule in 2008. The original researcher spent years publicly rejecting that interpretation.

What is the difference between an expert and a novice?

Experts organise knowledge around deep principles while novices organise it around surface features. In a classic sorting study physicists grouped problems by the underlying law required while beginners grouped them by visual appearance. Experts also perceive large meaningful patterns rapidly and retrieve relevant knowledge with much less conscious effort.

How do experts remember so much information?

They do not have larger memory capacity. They store information in chunks and templates built through years of exposure, so a single retrieved unit contains what would otherwise be many separate items. Chess masters lose most of their memory advantage when pieces are placed in configurations that could never occur in a real game.

Can experts be confidently wrong?

Yes and it is measurable. The Einstellung effect shows that finding a familiar good solution blocks the search for a better one, with eye tracking revealing that experts believe they are still searching when they are not. Long term forecasting studies also found average expert accuracy barely above chance.

Does the brain physically change with expertise?

Yes though the evidence has limits. A longitudinal study scanned 79 trainee London taxi drivers before and after training and found grey matter growth in the posterior hippocampus only among those who qualified. Many other expertise imaging studies are cross sectional with small samples and replicate poorly.