Introduction

Put twenty-five chess pieces on a board in a position taken from a real game. Show it to a grandmaster for five seconds. Take it away. The grandmaster will rebuild almost the entire board from memory, piece by piece, square by square, with very few errors. Show the same board to someone who learned the rules last month and they will manage four pieces, maybe five.

The obvious explanation is that grandmasters have extraordinary memories. The obvious explanation is wrong.

Now take the same pieces and scatter them at random across the board. Positions that could never arise in a real game. Show it for the same five seconds. The grandmaster's advantage largely disappears. Chess masters, it turns out, are stuck with the same brutally small working memory as everyone else [1]. What they have is not more storage. It is a better filing system.

That filing system is called chunking memory, and it is one of the most quietly important ideas in cognitive science. It explains how a mind that can hold only about four separate things at once manages to read sentences, diagnose diseases, play the piano, and reconstruct a chessboard. It also explains why the skill does not transfer, why brain training games disappoint, and why an expert's greatest strength sometimes becomes a blindfold.

This is the story of that idea. It runs from a telegraph office in 1899, through a Dutch psychologist timing chess players with a stopwatch, to a Carnegie Mellon laboratory where two researchers measured the pauses between a master's hand movements, and into laboratories today where the argument about what a chunk actually is remains unsettled.

Overhead view of a chessboard with glowing pieces and golden threads.

The Integer That Persecuted George Miller

In April 1955, a psychologist stood in front of the Eastern Psychological Association and opened his talk with a complaint. He said he had been persecuted by an integer. For seven years, he said, the number had followed him around, intruded on his private data, and assaulted him from the pages of the most public journals [2].

The psychologist was George Miller, and the number was seven.

The paper he published the following year, "The magical number seven plus or minus two", became one of the most cited works in the history of psychology. It is also one of the most misread. Almost every article written about chunking memory since has reduced it to a single claim: your short-term memory holds seven items. Miller never quite said that, and he spent much of the paper arguing something far more interesting.

He was actually describing two separate limits that happen to land near the same number. The first is the span of absolute judgment. Play someone a tone and ask them to name its pitch on a scale, and they can reliably distinguish only about six or seven levels before they start making mistakes. This limit is measured in information, roughly two and a half bits. The second is the span of immediate memory, the number of things you can hold in mind right now. That limit is not measured in information at all. It is measured in items.

Miller was explicit that the two limits are different in kind, and he treated their numerical similarity as a coincidence rather than a law. Absolute judgment is limited by the amount of information. Immediate memory is limited by the number of items. That distinction is the hinge of the whole paper, and almost nobody quotes it.

Which leads to the part that mattered most. If immediate memory counts items rather than information, then the size of the item is up for negotiation. Miller called the item a chunk, and he pointed out that recoding smaller units into bigger ones is the standard trick people use to beat the limit. A string of binary digits is hopeless. Group them into threes, translate each group into a single decimal digit, and suddenly the same information fits.

The chunk was born as a unit of convenience. It would take another two decades before anyone worked out how to measure one.

Then the Magic Number Turned Out to Be Four

Miller hedged. The field did not.

For decades, seven plus or minus two hardened from a rough observation into a textbook fact. Then in 2001, Nelson Cowan published a paper in Behavioral and Brain Sciences that quietly dismantled it [3]. His argument was not that Miller had been careless. It was that Miller's number had been inflated by the very trick Miller described.

Here is the problem. If you ask someone to remember a list, they will do two things automatically. They will rehearse it under their breath, refreshing the items before they fade. And they will group items together using knowledge they already have, turning three digits into a year or four letters into a word. Both of these inflate the score. Neither of them tells you how much raw capacity the system has.

Cowan went through the literature looking for studies that had blocked both escape routes. He found four kinds of evidence. Studies where information overload forced chunks down to single items. Studies that deliberately prevented recoding. Studies showing sharp performance discontinuities at a specific load. And indirect effects of the same limit showing up in other tasks.

The four converged on a single answer. When you strip out rehearsal and prior knowledge, the number of independent chunks a person can hold is about four, with a working range of three to five.

Visual memory told the same story from a different direction. In 1997, Steven Luck and Edward Vogel had shown people arrays of coloured bars and asked them to detect changes after a brief blank screen [4]. Performance held steady up to about four objects and then fell apart. What made the result striking was that each object could carry several features at once. Four bars meant four colours and four orientations, eight features, held without extra cost. The limit was not counting features. It was counting objects.

Alan Baddeley, writing about Miller's legacy, had already noted that the magic number was doing more cultural work than scientific work [5]. Cowan later summarised the revised position under a title that captured the mood, calling it the magical mystery four [6].

So what does this mean for anyone trying to learn something difficult? It means the bottleneck is far tighter than most study advice assumes. Four slots is not much. Any strategy that involves holding six unrelated new facts in mind while you work on a seventh is fighting the hardware. The only real way through is to make the units bigger, which is exactly what chunking does, and which is why the topic overlaps so closely with cognitive load in instructional design.

Four slots. That is the constraint every expert in every field has to work around. The clearest demonstration of how they do it came from a chessboard.

Estimated chunk capacity reported by four landmark studiesMiller 1956Luck 1997Cowan 2001Gobet 2004876543210Chunks held

The downward trend is not a sign that people got worse at remembering. It is a sign that experimenters got better at closing the loopholes.

From Amsterdam to Pittsburgh: Making a Chunk Visible

Long before anyone measured a chunk, a Dutch psychologist named Adriaan de Groot went looking for the source of chess genius in the most direct way possible. He asked strong players to think out loud.

De Groot's 1946 dissertation, later translated into English as Thought and Choice in Chess, set out to find what separated grandmasters from merely good players. The intuitive answer was calculation. Masters must see further ahead, consider more moves, search more branches of the tree.

They did not. De Groot found that grandmasters searched roughly the same number of possibilities as weaker players, sometimes fewer, and almost certainly not more. What they were extraordinarily good at was arriving at the right candidate move in the first place. The search was not deeper. The starting point was better.

Then he ran the experiment that mattered. He showed players a position from a real game for a few seconds and asked them to reconstruct it. Masters rebuilt these positions with near-perfect accuracy. Below master level, the ability dropped away sharply.

This was a genuine puzzle. If masters had generally superior visual memory, that would explain it, but there was no other evidence that chess players had unusual memories for anything else. Something specific was happening, and it was happening in the first few seconds of looking.

De Groot had found the phenomenon. He had not found the mechanism. That took another twenty-seven years, a stopwatch, and a pair of researchers willing to measure something almost nobody had thought to measure.

That measurement arrived in 1973, when William Chase and Herbert Simon published a paper in Cognitive Psychology called "Perception in chess" [1]. It is, by any reasonable measure, the most important experimental paper ever written about chunking memory, and its central method was almost absurdly simple.

They worked with three players. One master, one Class A player, and one beginner. Small sample, which is worth noting, but each player was tested across many trials on twenty positions taken from real games. Ten were middlegame positions with roughly twenty-four to twenty-six pieces on the board. Ten were endgame positions with twelve to fifteen. Eight further boards were built by scattering the same pieces at random.

There were two tasks. In the first, the target board stayed in full view and the player copied it onto a second board. Since the original was always available, any memory involved lasted only as long as a single glance. In the second task, the board was shown for five seconds and then removed, and the player rebuilt it from memory.

Here is the piece of method that changed everything. Chase and Simon watched the timing of the hand movements. Players did not place pieces at a steady rate. They placed several pieces in a rapid burst, then stopped, then placed another burst. The researchers proposed that the long gaps marked the boundaries between chunks. Everything placed inside a burst came out of memory together as a unit. Everything after a pause came from a different unit.

They set the boundary at two seconds.

Then they tested whether the boundary meant anything. If pieces within a burst really belonged to one meaningful pattern, they should be related to each other in chess terms. Chase and Simon scored five relationships between consecutive pieces: attack, defence, physical proximity on the board, shared colour, and shared piece type. Pieces placed inside the same burst shared these relationships far more often than pieces separated by a pause.

The pause was not noise. It was the seam between two units of knowledge.

That single methodological choice turned an invisible mental structure into something you could count with a stopwatch. Twenty-five years later, Fernand Gobet and Simon repeated the study with better equipment and found the two-second criterion held up, and that the chunks it revealed had genuine psychological reality [7].

But the most important result in the 1973 paper was not the pause. It was what happened when the pieces stopped making sense.

Two identical wooden chessboards with glowing clusters of pieces.

Scramble the Board and the Genius Disappears

This is the finding that almost no popular article about chunking memory bothers to mention, and it is the most revealing result in the entire field.

Chase and Simon built control boards by taking the same number of pieces and placing them at random. Same material. Same visual complexity. Same five seconds of viewing. The only thing removed was meaning.

The master's advantage collapsed [1]. On random boards, he reconstructed positions no better than the weaker players. The authors drew the conclusion plainly: masters appear to be constrained by exactly the same severe short-term memory limits as everyone else.

Think about what that rules out. It rules out superior raw storage. It rules out faster visual processing in any general sense. It rules out photographic recall, an idea that turns out to be shaky in almost every context where it has been tested and which has its own long history of disputed evidence. Whatever the master had, it was tied to the specific patterns that arise in real games, and it vanished the moment those patterns did.

From this, Chase and Simon reasoned their way to a number. If a master can hold the locations of twenty or more pieces but has room for only about five chunks, then each chunk must contain roughly four or five pieces bound into a single relational structure. A castled king position. A pawn chain. A familiar rook and knight configuration. Not twenty-five separate facts. Five patterns.

This is the core insight of chunking memory, stated as cleanly as it has ever been stated. Experts do not enlarge the container. They enlarge what fits inside each slot.

What does this mean outside chess? It means that when someone in your field seems to absorb a complex situation instantly, they are almost certainly not processing more raw detail than you. They are recognising fewer, larger things. A radiologist does not scan a scan pixel by pixel. An experienced developer does not read a function line by line. And crucially, that ability is welded to the domain. Scramble the pattern and the advantage goes with it, a point explored further in the psychology of expertise.

Which raises the obvious next question. How many of these patterns does an expert actually have?

Dense honeycomb with crystalline shapes and empty surrounding cells, warmly backlit.

Fifty Thousand Patterns and the Ten-Year Rule

The same year as "Perception in chess", Simon and Kevin Gilmartin published something less famous and arguably more audacious. They wrote a computer program to do the task [8].

The program was called MAPP, short for Memory-Aided Pattern Perceiver. It scanned a chessboard, recognised familiar configurations using a branching network of tests, stored pointers to those configurations in a limited short-term memory, and then decoded the pointers to rebuild the board. It was a working model of the chunking theory rather than a description of it, and it reproduced the human pattern reasonably well.

The interesting part was the extrapolation. By varying how many patterns the program knew and seeing how its performance changed, Simon and Gilmartin estimated how many patterns a human master would need. Their answer was between ten thousand and one hundred thousand chunks. The figure that entered the literature, and that Gobet and Simon would still be citing as the usual estimate two decades later [9], was around fifty thousand.

Chase and Simon liked to compare it to vocabulary. A well-read adult knows roughly fifty thousand words and recognises almost all of them instantly without conscious effort. A chess master, on this account, has a comparable vocabulary made of board patterns instead of words.

Simon kept pulling at the thread. In 1974 he published a short paper in Science with the blunt title "How Big Is a Chunk?" [10]. He tested memory span for material of increasing size and found something awkward for the magic seven. Span for one-syllable words was around seven. For longer words it fell. For short phrases it fell further, toward three. The bigger the chunk, the fewer you could hold. Capacity in chunks was not a fixed seven at all.

There was one more consequence. If mastery requires tens of thousands of patterns, and patterns are acquired through exposure, then mastery requires time. Simon and Chase estimated that nobody reaches grandmaster strength on less than roughly a decade of serious work, an observation that later grew into the deliberate practice research programme [11]. That programme has since been sharply contested. A large meta-analysis found that accumulated practice accounted for around a quarter of performance variance in games and less in other domains [12], and studies of chess players specifically found that starting age and practice interact in ways the simple story missed [13]. The ten-year figure survives as a rough regularity, not a law.

The chunking theory was, at this point, in excellent shape. It had a mechanism, a measurement method, a computer model, and a number. Then someone rechecked the control condition.

1899
Bryan and Harter find plateaus in telegraphy learning
1946
De Groot times chess masters reconstructing real positions
1956
Miller names the chunk and the magical seven
1973
Chase and Simon scramble the chessboard
1974
Simon asks how big a chunk really is
1980
A student reaches seventy-nine digits by recoding
1996
Gobet and Simon propose template theory
2001
Cowan revises the capacity limit down to four
2019
Chunking splits into compression and reconstruction
2026
Synaptic model keeps four clusters active at once

The earliest entry on that list deserves a moment. In 1899, William Bryan and Noble Harter studied people learning telegraphy and found that progress came in steps rather than a smooth curve [14]. Learners plateaued, then jumped. Their explanation was that operators were building a hierarchy of habits, moving from letters to words to phrases as single units. That is chunking, described more than half a century before it had a name.

The Small Anomaly That Broke a Big Theory

Good theories die from small numbers.

Chunking theory made a sharp prediction about random boards. If the master's advantage comes entirely from recognising game patterns, and random boards contain no game patterns, then masters should perform exactly like novices on random boards. Chase and Simon's data appeared to show precisely that.

But their sample was three players. When Gobet and Simon pooled the random-position results from more than a dozen studies, a small, stubborn effect appeared [9]. Masters were slightly but reliably better on random boards. Not dramatically. Not the way they were better on real positions. But consistently, and not by chance. A companion study confirmed that recall of rapidly presented random positions also varied with skill [15].

By the letter of the original theory, that should not happen. Zero patterns should mean zero advantage.

There was a second problem, and it was worse. Masters can play blindfold chess against several opponents at once, holding multiple boards in mind simultaneously. Chunking theory said short-term memory holds around seven pointers. Seven pointers cannot hold four chessboards. The arithmetic simply does not work.

A third crack came from timing. Gobet and Simon later showed that masters extract a great deal from a position in one second and continue to gain from longer viewing up to a minute [16]. Information seemed to be going into long-term storage far faster than a short-term buffer model allowed.

Not everyone accepted the framework at all. Alexandre Linhares and Anna Freitas argued that experts recognise positions by deep strategic similarity rather than surface pattern, and that the chunk as measured by Chase and Simon was the wrong unit entirely [17]. Gobet defended the two-second criterion and the pattern account, and the exchange remains one of the more pointed disagreements in the field.

The response from Gobet and Simon was not to abandon chunking. It was to give the chunks somewhere bigger to live.

Templates: How Experts Cheat the Bottleneck

In 1996, Gobet and Simon proposed template theory [18]. The idea is best understood as a form filled in rather than a photograph taken.

A template is a large, familiar structure stored in long-term memory. It has a stable core, the squares and pieces that are always the same in this type of position, and a set of empty slots for the details that vary. When a master looks at a board, they do not encode thirty independent facts. They recognise a template, then fill its slots with the specifics of this particular game. Slot filling takes seconds, not minutes, which means the information lands in long-term memory almost immediately instead of queuing in a four-slot buffer.

That solves all three problems at once. Blindfold simultaneous chess becomes possible because each board is a template in long-term memory rather than a set of pointers in short-term memory. Rapid encoding becomes possible because a slot is faster to fill than a structure is to build. And the random-position anomaly becomes explicable because a master with tens of thousands of overlapping patterns will, purely by chance, find a few small familiar fragments even in a scrambled board.

Abstract grid of empty slots with glowing amber tokens on cream paper.

That last point was not just hand-waving. The computational model built on template theory, called CHREST, reproduced the small random-position effect as an emergent consequence of having more and larger stored patterns [19]. CHREST descends from an older program, EPAM, developed by Edward Feigenbaum and Simon, which learned by growing a branching network of discrimination tests [20]. Simon's original chunking estimate of fifty thousand applied to small chunks. Template theory implies a far larger store, since templates are bigger and more numerous than the four-piece fragments of the 1973 model, pushing the plausible figure into the hundreds of thousands.

Then came a measurement that sharpened the picture considerably. In 2004, Gobet and Gary Clarkson re-ran the copy and recall tasks using both physical boards and computer displays, with a within-subject design so the same players faced both conditions [21]. Two results stood out. In most cases, no more than three chunks were replaced during recall. And on game positions in the computer condition, masters replaced chunks containing up to fifteen pieces.

Three chunks. Fifteen pieces each. That is the entire trick of chunking memory in two numbers. The slots stayed brutally few. The contents grew enormous.

By 1998, Gobet had formally compared the competing accounts of expert memory and concluded that while chunking theory fits most of the data, template theory fits it best [22].

TheoryCore claimStrongest evidenceMain difficulty
Chunking theory (1973)Experts store tens of thousands of small patterns and hold pointers in short-term memoryRandom-position collapse and the two-second pause boundaryCannot explain blindfold simultaneous play or the small random-position effect
SEEK theory (1985)Expertise is mainly search and evaluation knowledge rather than perceptionAccounts for high-level planning in strong playersPoor fit to rapid recall and very short presentation times
Long-term working memory (1995)Experts build retrieval structures that extend working memory into long-term storageTrained digit-span cases and resistance to interferenceLess specific about the structure of chess knowledge itself
Template theory (1996)Chunks plus large schematic templates with fillable slotsBest fit to random and distorted positions plus interference dataRequires more theoretical machinery than its rivals

Long-term working memory, the third entry in that table, came from a different direction entirely. And it came from an experiment on a single undergraduate.

The Student Who Held Seventy-Nine Digits

In 1978, a student at Carnegie Mellon signed up for a psychology experiment. He is known in the literature as SF. His starting digit span was seven, which is to say, entirely ordinary.

Over roughly twenty months and more than two hundred and thirty hours of practice, one hour a day, several days a week, Anders Ericsson, Chase and Steve Faloon tracked what happened [23]. By the end of the study, SF could hear a list of seventy-nine random digits read at one per second and repeat it back correctly.

Seventy-nine. From seven.

The method was not mysterious. SF was a competitive distance runner, and he began mapping digit groups onto running times. A sequence like 3492 became three minutes forty-nine point two seconds, close to a world record mile. Groups that would not fit as race times became ages or dates. He built a hierarchy on top: groups of three or four digits combined into supergroups, supergroups into a retrieval structure that let him unload the whole list in order.

He was not expanding his memory. He was manufacturing chunks out of knowledge he already had.

And this is the part that matters most, the part almost every article about chunking memory leaves out. Throughout the entire training period, SF's memory span for letters stayed at about six.

Not seventy-nine. Six. Exactly where it started.

The skill did not transfer at all. It was welded to digits, because the encoding scheme was made of running times, and running times are made of digits. Take away the material the chunks were built from and the extraordinary performance evaporated, precisely as it did when Chase and Simon scrambled the chessboard.

SF continued past the published study and eventually reached around eighty-two digits over two hundred and sixty-four hours [24]. A later participant, known as DD, trained under the same regime and climbed from a span of eight to sixty-eight digits over two hundred and eighty-six hours, ultimately reaching one hundred and six before stopping in 1985 [25]. When researchers retested DD after three decades away from the task, the story of what survived and what faded became a small case study in its own right.

Out of this work came the theory of long-term working memory, developed by Ericsson and Walter Kintsch [26]. The claim is that skilled performers build retrieval structures in long-term memory that behave like an extension of working memory within their domain. It is a different route to the same destination as template theory: the bottleneck is not widened, it is bypassed.

What does this mean for someone studying? It means the useful question is never how to hold more items. It is what the learner already knows well enough to hang new material on. SF had running times. A medical student has anatomy. A musician has scales. Chunks are built from existing knowledge, which is why the first hours in any new field feel so much harder than the hundredth.

Tall narrow spiral staircase with indigo geometric blocks ascending upward.

What the Brain Does When It Chunks

For most of this story, the chunk was an inference. You could see its shadow in a pause between hand movements or a jump in a learning curve, but nobody had watched one form.

Then imaging arrived, and the first clear result was the opposite of what everyone expected.

In 2003, Daniel Bor, John Duncan, Richard Wiseman and Adrian Owen scanned people while they held sequences of digits in mind [27]. Some sequences had hidden structure that made them easy to group. Others were random. Structured sequences were easier, produced better recall, and by any sensible definition placed less demand on memory.

They produced more prefrontal activity, not less.

Specifically, lateral prefrontal cortex, the region running along the outer front surface of the brain that handles planning and the manipulation of information, lit up more strongly during the encoding of structured material. A follow-up study confirmed that both the upper and lower parts of that region are recruited when material is being recoded [28], and later work showed the same prefrontal and parietal network handling both mnemonic and mathematical recoding [29].

The interpretation is that chunking is not a saving. It is a trade. The brain spends effort at the moment of encoding, running a search for structure, in exchange for a lighter load afterwards. Bor and Anil Seth later folded this into a broader argument that this same prefrontal-parietal machinery underpins the conscious, effortful search for pattern [30].

Chess players have been scanned too. Guillermo Campitelli, Gobet and colleagues looked specifically for where chess chunks live in the brains of players of different strengths [31]. And in 2012, Alessandro Guida and colleagues proposed a two-stage account that reconciles a confusing split in the imaging literature [32]. Early in skill acquisition, activation in task-relevant regions decreases as processing becomes efficient. Later, once knowledge structures and templates are in place, the brain reorganises functionally, recruiting long-term memory regions to do working-memory work. Two apparently contradictory findings, one developmental sequence.

The mechanism question is being attacked from another angle entirely. Building on the idea that working memory can be held in the temporary strength of synapses rather than in continuous firing [33], a recent model by Weishun Zhong, Mikhail Katkov and Misha Tsodyks proposes a specific circuit for chunking [34]. Dedicated chunking clusters group the individual item representations. The network suppresses and reactivates these groups in turn, so that all items in several chunks can be recovered while no more than about four clusters are ever active at the same moment. The four-slot limit is never violated. It is worked around. The paper is currently a reviewed preprint rather than a finished journal article, so the finding should be treated as promising rather than settled.

And chunking is not only a memory phenomenon. It is deeply motor. Ann Graybiel argued in 1998 that the basal ganglia, a set of deep brain structures involved in habit and action selection, package action sequences into single performance units [35]. Recordings later revealed start and stop signals emerging in these circuits as sequences become automatic [36]. Humans learning visuomotor sequences spontaneously carve them into small groups of three or four elements [37], and imaging during motor chunking shows the putamen and frontoparietal cortex taking on different roles as the sequence consolidates [38].

This is why a pianist stops thinking in notes and starts thinking in phrases, and why the process by which habits become automatic looks so much like the process by which chess positions become patterns. The same compression logic, running in a different part of the brain.

Abstract network of glowing nodes in luminous bubbles on dark background.

Compression or Reconstruction: An Argument Still Running

Here is the honest position. After seventy years, researchers still disagree about what chunking actually does to the contents of your mind.

There are two answers, and they are genuinely different.

The first is compression. On this account, chunking recodes the input into a smaller representation that occupies less space. Three words become one unit, and that unit takes up one slot instead of three. The evidence is real. Fabien Mathy and Jacob Feldman showed that the compressibility of a sequence predicts how much of it people can recall, and that the underlying limit in truly incompressible units sits closer to three or four than seven [39]. Timothy Brady and colleagues found the same logic operating in visual memory, where statistical regularities across an array allow more to be stored [40]. Matthew Nassar, Julie Helmers and Michael Frank went further, modelling chunking as rational lossy compression, with capacity bought at the price of precision [41]. Later work tied chunk formation directly to compression in immediate memory [42].

Translucent glass sheets and a kintsugi-repaired ceramic bowl.

The second answer is redintegration, an old word for reconstruction. On this account, chunks never enter short-term memory at all. They sit in long-term memory, and when the short-term trace has degraded into something fuzzy and partial, long-term knowledge is used to rebuild it. Nothing was compressed. Something was repaired.

The two accounts make different predictions, and in 2020 Dennis Norris, Kristjan Kalm and Jane Hall built experiments to separate them [43]. The result was not a clean victory for either side. Small chunks of two words behaved like redintegration. Larger chunks of three words behaved like data compression. And recoding material into chunks carried a measurable cost of its own. Norris and Kalm pursued the compression side further the following year [44].

Pulling in the other direction, Mirko Thalmann, Alessandra Souza and Klaus Oberauer had already argued for a strong version of the capacity-freeing view [45]. In their experiments, a chunk reduced load by pulling a single compact representation out of long-term memory to replace the individual elements. The telling detail was that chunking one part of a list improved memory for other, unchunked items presented alongside it. Capacity had genuinely been freed, not just reconstructed.

Underneath all of this sits an uncomfortable methodological problem. Amy Gilchrist pointed out that the field has never agreed on how to measure a chunk, and that different measurement choices produce different conclusions about the same data [46]. Gary Jones has argued that chunking is underused as an explanation in developmental psychology, where it could account for changes usually attributed to raw capacity growth [47]. A large collaborative effort to define benchmarks for working memory models exists partly because these disagreements were becoming difficult to adjudicate [48].

The most defensible reading of the current evidence is that both mechanisms are real and which one dominates depends on chunk size. That is a less satisfying answer than a single mechanism. It is also, at the moment, the honest one.

When Good Chunks Block Better Ones

Every account so far has treated chunking as an advantage. It is time to look at the bill.

The first cost is the one SF paid. Chunks are built from domain knowledge, so they only work inside that domain. This is not a minor caveat. It is the reason an entire commercial category has struggled to justify itself. A major review of brain-training programmes by Daniel Simons and colleagues found that practising a task reliably improves that task, produces some improvement on closely similar tasks, and offers very little evidence of transfer to everyday cognitive performance [49]. Meta-analyses of working memory training reached the same conclusion, finding short-lived near transfer and no convincing far transfer [50], a result that held up on re-examination [51]. Even chess instruction, the obvious candidate given everything above, shows only small effects on academic and cognitive skills that shrink further once study quality is accounted for [52].

The second cost is stranger and more interesting. Familiar patterns do not just help. Sometimes they actively block.

In 2008, Merim Bilalić, Peter McLeod and Gobet designed a chess problem with two solutions [53]. One was a well-known checkmating motif that worked but took five moves. The other was a shorter, better solution of three moves. Players who knew the famous motif found it immediately and then stopped looking. When told a better solution existed and asked to keep searching, they reported that they were still looking.

Their eyes said otherwise. Eye-tracking showed that gaze kept returning to the squares relevant to the first solution, even while players insisted they had moved on. The familiar chunk had captured attention below the level of awareness. This is the Einstellung effect, and its cost was measurable: it dropped skilled players to the performance level of players rated several classes below them. The same team quantified how the effect scales with expertise across a wider set of problems [54].

One detail is worth holding on to. The very strongest players tended to escape the trap on the problems tested. Enough patterns, and the first one no longer monopolises the search. The blindness sits in the middle of the skill curve, not at the top.

There is a broader lesson here about how expertise fails. It rarely fails by producing nothing. It fails by producing something good, quickly, and then declining to look further. Anyone who has watched an experienced clinician anchor on a first diagnosis, or an experienced engineer reach instantly for the architecture that worked last time, has seen a chunk doing exactly what it was built to do, in a situation where it should not have.

What does this mean in practice? It means deliberately generating a second candidate before evaluating the first is not a productivity ritual. It is a countermeasure against a documented perceptual bias, and it matters most for people who are good but not exceptional at what they do.

Well-worn path through tall grass leading to a bright horizon.

Chunks Everywhere: Fingers, Phrases and Diagnoses

The phone number example has been doing the work in explanations of chunking memory for seventy years. It undersells the idea badly.

Start with language. Fluent speakers do not assemble sentences word by word from grammatical rules. They deploy stored multi-word units. Inbal Arnon and Neal Snider showed that people process frequent four-word phrases faster than infrequent ones, even after controlling for the frequency of the individual words [55]. The phrase itself is the unit. This is why a learner who has memorised a thousand words still sounds laboured while a child who has absorbed a thousand phrases sounds natural.

Language chunking starts absurdly early. Jenny Saffran, Richard Aslin and Elissa Newport played eight-month-old infants a two-minute stream of nonsense syllables with no pauses and no intonation cues, in which the only signal was how often one syllable followed another [56]. The infants extracted the word boundaries. Before they can speak a word, they are already computing which sounds belong together.

Then medicine. Experienced clinicians do not reason from individual symptoms to a diagnosis by exhaustive elimination. They recognise illness scripts, structured packages of typical presentation, timing and context that behave exactly like templates with fillable slots [57]. A junior doctor holds a list of findings. A senior one holds a pattern. This is the same architecture that lets the brain build categories from clinical patterns, and it carries the same Einstellung risk when the presentation is atypical.

And then the telegraph operators from 1899, who turned out to be describing everything above before any of the vocabulary existed. Their plateaus were not fatigue. They were the periods spent building the next level of the hierarchy, from letters to words to phrases, during which visible performance stalls while structure forms underneath.

Four slots, filled with progressively larger things. That is the whole architecture, running in the hands of a pianist, the ear of an infant, the eye of a radiologist and the fingers of a telegraph operator.

For anyone trying to learn deliberately, the practical consequence is specific. Material that has no internal structure cannot be chunked, so imposing arbitrary groupings on it will not help and may cost effort for nothing. Material that does have structure should be studied in a way that lets the structure become visible, which usually means retrieving it rather than rereading it [58]. And the size of the unit should grow over time. Practising the same four-note fragment forever builds a very good four-note chunk and nothing above it. Deliberate practice research has been criticised on many points, but the observation that practice must keep raising the size of the unit has held up reasonably well [59].

Grand piano keyboard with warm golden keys and cool shadows.

Conclusion

The story of chunking memory is really the story of a number getting smaller and a unit getting bigger.

Miller offered seven and admitted he was half joking about it. Cowan, with better controls, brought it down to four. Gobet and Clarkson, watching masters rebuild chessboards, found that in practice no more than three chunks were being moved at once. Every improvement in method has tightened the bottleneck.

And in the same seventy years, the chunk itself grew from a three-digit group to a fifteen-piece chess configuration to a template with fillable slots to a retrieval structure that behaves like an annexe of working memory. The container never expanded. The contents did.

What makes this more than a technical footnote is what it implies about talent. The grandmaster who rebuilds a board in five seconds is not operating different hardware. Scatter the pieces and that becomes obvious within a single trial. The student who recited seventy-nine digits could still only manage six letters. The expertise was never in the machinery. It was in the tens of thousands of hours of pattern that had been loaded into it, and it stayed exactly where it was put.

That is the encouraging half of the finding and the sobering half at once. Nobody is born with more slots. But nobody gets to borrow someone else's chunks either.

There is one loose thread, and it is the honest place to end. Researchers still cannot agree whether a chunk compresses information inside short-term memory or sits in long-term memory rebuilding what short-term memory dropped. The current evidence points at both, depending on how big the chunk is. Seventy years after Miller stood up and complained about being persecuted by an integer, the integer has been corrected, the mechanism has been imaged, the model has been coded, and the most basic question about what a chunk actually is remains open.

Which is roughly how it should be. The board looks simple until you count the pieces.

Frequently Asked Questions

How many chunks can working memory actually hold?

Modern estimates put the limit at about four, with a working range of three to five, once rehearsal and prior knowledge are experimentally blocked. Miller's familiar figure of seven plus or minus two was inflated because participants were quietly recoding and rehearsing. Studies of expert chess recall suggest the practical figure may be closer to three.

Does chunking make you better at remembering everything?

No. Chunks are built from knowledge you already have in a specific domain, so the benefit stays inside that domain. The most famous demonstration involved a student who trained his digit span from seven to seventy-nine while his memory span for letters never moved beyond about six.

Why can chess masters memorise a board so quickly?

They recognise familiar configurations rather than individual pieces, so twenty-five pieces become a handful of known patterns. The proof is the control condition. When the same pieces are scattered into positions that could never occur in a real game, the master's advantage largely disappears.

Is chunking the same thing as using a mnemonic?

They overlap but are not identical. A mnemonic is a deliberate technique for imposing structure on material, often using imagery or rhyme. Chunking is the underlying process of binding elements into single units, and it happens automatically through exposure as well as deliberately through strategy.

Does chunking happen outside memory tasks?

Yes. The same grouping logic appears in motor learning, where the basal ganglia package action sequences into single units, in language, where frequent multi-word phrases are processed as wholes, and in clinical reasoning, where experienced doctors recognise structured illness patterns rather than isolated symptoms.