Introduction
Look at a die. You know it landed on four. You did not count the pips, you did not say the numbers in your head, and you did not decide to do anything at all. The four simply arrived. Now look at a scattered handful of coins on a table. Eight of them, maybe nine. Suddenly you are working. Your eyes move, your attention hops from coin to coin, and somewhere behind your forehead a slow serial machine grinds into gear.
That gap between the effortless four and the laborious eight has a name. Subitizing, from the Latin subitus, meaning sudden. It is the ability to apprehend the number of a small set rapidly, accurately, and without any sense of effort [2]. It is also one of the most stubbornly replicated findings in cognitive science, first measured in 1871 by an economist throwing beans into a box [1], and still the subject of an unresolved argument among vision scientists today.
Most of what has been written about subitizing is aimed at kindergarten teachers. That is a shame, because the interesting question is not how to teach it to a five year old. The interesting question is what it reveals about the architecture of an adult mind. Why does the number four keep appearing in unrelated corners of psychology? Why do bees, crows, chimpanzees and humans all appear to share some version of this ability? And why, after 150 years, can researchers still not agree on whether subitizing is a special mechanism at all?
This is the story of that argument.

A Victorian Economist and a Box of Beans
William Stanley Jevons was not trying to found a field. He was a logician and economist at Owens College in Manchester, and in 1871 he was curious about the limits of his own mind. So he built an experiment out of household objects.
He took a shallow paper box and a supply of black beans. He would scoop up a random handful, throw it into the box without looking, glance at the result, and immediately write down how many beans he believed were there. Then he counted them properly. He repeated this 1,027 times.
The results were unambiguous. For three beans and four beans, he was never wrong. Not once. At five, errors crept in. At six, he was correct on 120 of 147 attempts. At seven, 113 of 156. By the time he reached ten or twelve beans, his guesses had become genuinely unreliable. Jevons fitted a rough function to his own error rate and concluded that his mind held a limit somewhere near four and a half [1].
He published it in Nature, in a short paper that reads more like a diary entry than a scientific report. And then, for the better part of eighty years, almost nobody did anything with it.
The revival came in 1949, at Mount Holyoke College. E. L. Kaufman, M. W. Lord, T. W. Reese and J. Volkmann ran a far more controlled version of the same question. Rather than beans in a box, they flashed arrays of dots on a screen for a fixed 200 milliseconds, too brief for the eyes to move and count. Set sizes ran from a single dot up to more than two hundred. What they found was that performance was essentially perfect up to about five items, then degraded sharply.
More importantly, they gave the phenomenon a name. Arguing that the process for small numbers was neither counting nor estimating but something distinct, they coined the verb "subitize" from the Latin adjective subitus [2]. They separated three processes explicitly: subitizing for small sets, counting for larger sets, and estimating for sets too large or too brief to count.
This matters for a reason that has nothing to do with history. A remarkable number of teaching websites, parenting blogs and educational resources state that the term was coined by Jean Piaget. It was not. Piaget wrote extensively about children's number concepts, and his influence on early mathematics education is enormous, but the word belongs to Kaufman and his three colleagues in 1949. The misattribution has been copied so many times that it now appears near the top of search results for the term. If a claim about subitizing seems to trace back to Piaget, it is worth checking the original.
A year earlier, Irving Saltzman and Wendell Garner had made a methodological move that turned out to be decisive. They proposed using reaction time, rather than accuracy alone, as the measure of how many items a mind can grasp at once [3]. Accuracy tells you when a system breaks. Timing tells you how it works while it is still succeeding. That shift set up everything that followed.

The Signature Hidden in Milliseconds
Here is the finding that anchors everything else in this article, and it is worth stating precisely.
When people enumerate small sets, adding one more item costs them very little time. Somewhere between 40 and 100 milliseconds per item. Cross a boundary at around four, and the cost per item jumps by a factor of roughly five, to somewhere between 250 and 350 milliseconds. Errors follow the same pattern. Near zero below the boundary, rising steadily above it. Lana Trick and Zenon Pylyshyn laid out these numbers in a 1994 review that remains the standard reference for the effect [4].
Plot reaction time against set size and you do not get a straight line. You get two lines meeting at an elbow. That elbow is the subitizing limit.
The obvious objection is that the effect is about eye movements. Perhaps small sets are simply small enough to fit in one fixation, and larger sets require the eyes to travel. Janette Atkinson, Fergus Campbell and Marcus Francis tested this in 1976 with a method that is still elegant half a century later.
They dark adapted their observers, then fired a photographic flashgun at a line of white discs. The flash burned a bright afterimage onto the retina. Crucially, an afterimage moves with the eye. Look left and it goes left with you. There is no way to scan it, and no way to count by fixating each element in turn. Observers were asked how many discs they had seen, both ten seconds and sixty seconds after the flash.
Even with a full minute of viewing time, the limit held at four. Below that, no errors. Above it, consistent mistakes that a minute of staring could not fix. The authors titled the paper "The Magic Number 4 ± 0" [5]. They also found something that gets quoted less often. Squeeze the discs closer than about five hundredths of a degree of visual angle apart and the limit collapses to two. Spacing matters.
Tony Simon and Sandeep Vaishnavi replicated the afterimage approach twenty years later with three adult participants, an unusually small sample even for psychophysics. Their result was clean. Enumeration of afterimages was effectively error free for two, three and four discs, and produced error rates of twenty to thirty percent from five discs upward [6]. They read this as evidence that counting depends on shifts of attention while subitizing does not.

Not everyone reads the elbow the same way. George Mandler and Billie Shebo argued in 1982 that subitizing is not a numerical mechanism at all but a pattern recognition one [7]. Two items make a line. Three make a triangle. Four make a quadrilateral. Beyond four, the number of possible configurations explodes and no single shape can be matched. On this view, the mind is not counting quickly. It is recognising a familiar arrangement, which is a different operation entirely and connects to the broader machinery of visual pattern recognition.
The pattern account explains a lot, but not everything. Subitizing survives when the items are placed in random positions that form no recognisable shape, which a pure template matching account struggles to accommodate.
One more piece of the classic literature deserves mention. Michelene Chi and David Klahr measured the elbow in children and adults in 1975 and found that its position moves [8]. Young children show a shallower subitizing range, roughly three, and a steeper counting slope. Adults reach four or occasionally five. So the limit is not a fixed constant carved into the visual system. It develops.
What does this mean in practice? It means the four item boundary is a soft, context dependent frontier rather than a hard wall. Change the spacing, change the modality, change the age of the observer, and the number shifts. Anyone quoting "the subitizing limit is four" as a bare fact is compressing a much more interesting picture.
Two Systems, or Possibly Three
If subitizing handles one to four, what handles the rest?
The standard answer is a second system, older and cruder, called the Approximate Number System. It does not deliver exact answers. It delivers a fuzzy sense of quantity that gets fuzzier as numbers grow. Its behaviour follows a rule borrowed from nineteenth century psychophysics known as Weber's law, which says that discriminability depends on the ratio between two quantities rather than the absolute difference between them.
Telling ten dots from twenty is easy. Telling ninety from a hundred is hard. The difference is ten in both cases, but the ratio is not.
Robert Moyer and Thomas Landauer found the first clean signature of this in 1967, when they showed that the time it takes to judge which of two digits is larger depends on how close the digits are [10]. Comparing two and nine is fast. Comparing eight and nine is slow. That effect, the numerical distance effect, appears wherever the approximate system is doing the work.
Lisa Feigenson, Stanislas Dehaene and Elizabeth Spelke pulled the two system picture together in an influential 2004 review [11]. One system for small exact quantities, limited to about three or four. One system for large approximate quantities, limited by ratio. Both present early in life, both present in other species.
The approximate system sharpens with age, and the numbers are strikingly regular. Justin Halberda and Feigenson measured the Weber fraction, which is a compact way of describing how fine a ratio someone can reliably discriminate, across ages three to six and in adults [13]. Three year olds could manage a three to two ratio. Four year olds, four to three. Five year olds, five to four. Six year olds, seven to six. Adults, roughly ten to nine.
Lower bars mean finer discrimination. The curve does not flatten out in childhood, which is one reason researchers treat approximate number acuity as a developing skill rather than a fixed endowment. In a separate study of fourteen year olds, Halberda and colleagues found that individual differences in this acuity correlated with mathematics achievement stretching back to kindergarten [14]. Correlation, not causation, and the authors were careful to say so.
Then the picture got more complicated.

David Burr and Giovanni Anobile at the University of Florence, working with Guido Marco Cicchini in Pisa, have spent more than a decade arguing that two regimes are not enough. Their proposal is three. Subitizing for very few items. Estimation obeying Weber's law for sparse arrays. And a third regime for dense arrays, where the visual system stops tracking individual objects and starts reading the display as a texture, following a square root law rather than Weber's law [16].
The dividing line falls at roughly two items per square degree of visual angle. Below that density, the mind sees objects. Above it, the mind sees stuff [15]. The same group has shown that people register numerosity spontaneously, without being asked to attend to it, which is the sort of thing normally reserved for basic properties like colour or motion [25].
Antonella Pomè and colleagues found a reaction time signature that fits this three way split. Timing is flat and fast in the subitizing range, rises sharply once the set exceeds four, and then falls again for very dense displays. The point where the reaction time curve turns back down coincides with the transition from Weber's law to the square root law [17]. Sixteen participants, so this is not a large study, but the pattern is coherent.
What does this mean for someone reading a chart or scanning a spreadsheet? Roughly this. Your visual system is not applying one strategy to quantity. It is silently switching between at least two and possibly three, and the switch depends on how crowded the display is rather than on any decision you make.
The Argument That Will Not Die
Now the honest part.
Nobody disputes that the elbow exists. What people dispute, and have disputed for thirty years without resolution, is what causes it.
One camp holds that subitizing depends on a dedicated mechanism that operates before attention arrives. The most developed version comes from Trick and Pylyshyn, who proposed that the visual system carries a small pool of pointers, roughly four or five of them, that latch onto objects and keep track of them individually [4]. Pylyshyn called these pointers FINSTs, short for fingers of instantiation, and the same pool had already been invoked to explain how people can track several moving targets at once in a field of identical distractors [20]. On this account, subitizing is not counting at all. It is reading off how many pointers happen to be occupied. The pair had argued a year earlier that enumeration tasks are among the cleanest windows onto how spatial attention is allocated [9].
The strongest evidence for a dedicated mechanism comes from neurology. In 1994, Dehaene and Laurent Cohen studied five patients with simultanagnosia, a rare disorder in which someone can recognise individual objects perfectly well but cannot take in a visual scene as a whole. Asked to enumerate more than three items, these patients failed almost completely, either skipping objects or counting the same one twice. Asked to enumerate one, two or three items, they were excellent [37]. Their damage was bilateral and parietal, in tissue associated with shifting attention through space.
If a person who has lost the ability to move attention around a scene retains the ability to grasp small numbers, that is hard to explain unless small numbers are handled by something else.
The other camp is not persuaded, and it has data.
Petra Vetter, Brian Butterworth and Bahador Bahrami gave fourteen participants a demanding secondary task while asking them to enumerate. If subitizing were genuinely preattentive, loading attention elsewhere should leave it untouched. It did not. Accuracy for small numerosities degraded substantially. The paper's title says it plainly: evidence against a preattentive subitizing mechanism [21].
Burr, Marco Turi and Anobile pushed further. Under high attentional load, the precision of judgments in the subitizing range fell until it matched the precision of the estimation range, while estimation itself was barely affected [22]. Their conclusion inverts the classical picture. It is subitizing, not estimation, that is the attention hungry process. Henry Railo and colleagues in Finland reported converging results the same year [23], brain imaging by Daniel Ansari and colleagues found the small number range recruiting regions associated with attention rather than bypassing them [41], and a later study found the interference persists even when the competing task is in a different sense modality altogether [24].
Then there is a finding that is harder to fit into either box. Wei Liu and colleagues, working with Cicchini, showed participants two intermingled groups of dots distinguished only by colour and asked them to enumerate one group. Subitizing collapsed. The advantage that appears reliably for a single small set simply vanished when a second set shared the display. Estimation, meanwhile, coped fine [26]. In 2024 the same group found that subitizing survives when two sets are compared one after another, but not when they are compared at the same instant, where a ratio effect appears even for numbers as small as two and three [27].
Sitting between the camps is a study by Susannah Revkin, Manuela Piazza, Véronique Izard, Cohen and Dehaene, which tested whether the subitizing advantage is really about small numbers or about small ratios. They compared performance on one to eight items against performance on ten to eighty items scaled to preserve the same ratios. The subitizing advantage did not transfer to the large scaled range, which they read as evidence that something special really is happening below four [18].
And in 2020, Samuel Cheyette and Steven Piantadosi took a different approach entirely. Rather than arguing for one mechanism or two, they built a single model in which a mind allocates limited information to representing quantity, and showed that the elbow falls out of the mathematics on its own [19]. No dedicated small number module required. Four preregistered experiments with a hundred participants each.
So where does that leave things?
Genuinely unresolved. The attention load evidence is strong and has been replicated by independent groups. The neuropsychological dissociation is also strong and is not easily explained away. It is possible that both are describing real properties of a system nobody has yet characterised correctly. What can be said with confidence is that the sentence "subitizing is a preattentive parallel process" is stated as settled fact on a great many websites, and it is not settled.

Neurons That Keep Count
While psychologists argued about mechanisms, neurophysiologists went looking for the hardware.
Andreas Nieder, then working with Earl Miller at MIT, found it. In 2002 they recorded from single neurons in the lateral prefrontal cortex of rhesus monkeys trained to judge whether two displays contained the same number of items. Some neurons fired most strongly for displays of one item. Others preferred two, or three, or four, or five. Each neuron had a preferred numerosity and responded less as the display moved away from that preference in either direction [28].
These are tuning curves, the same kind of overlapping filters the visual system uses for orientation or colour [30]. And the curves have a revealing property. They get broader as numbers get larger, which is exactly what would produce Weber's law at the behavioural level. The distance effect and the size effect, both measured in humans decades earlier, turn out to be readable in the firing of individual cells.
Nieder and colleagues later mapped a parieto-frontal network for number, with the intraparietal sulcus as a central node [29]. The intraparietal sulcus is a groove running along the upper side of each parietal lobe, above and behind the ear. Human brain imaging had already implicated the same region. Piazza and colleagues used a technique called adaptation, in which repeated presentation of one numerosity dampens the response, to demonstrate numerosity tuning curves in the human intraparietal sulcus that closely resemble the monkey data [12].

Then the story took a turn that nobody expected.
Helen Ditz and Nieder went looking for number neurons in crows. Birds do not have a layered cortex. Their forebrain is organised on an entirely different plan, and the last common ancestor of crows and primates lived roughly three hundred million years ago. If number neurons appeared in both lineages, they appeared independently.
They did. Recording from a crow forebrain region called the nidopallium caudolaterale, the researchers found that more than twenty percent of randomly sampled neurons were significantly modulated by numerosity, about ninety eight cells out of four hundred and ninety nine. Only eight percent responded to both numerosity and the type of stimulus, meaning most of these cells cared about how many and not about what [31].
A follow up went further. Recording from crows that had never been trained on any numerical task, Lysann Wagener and colleagues found neurons that encoded numerosity spontaneously [32]. The tuning was there before the training. And in 2021, the group reported crow neurons that represent empty sets, treating zero as the low end of a numerical continuum rather than as nothing at all [33].
The human evidence arrived in 2018. Esther Kutter, working with Florian Mormann and Nieder, recorded from single neurons in the medial temporal lobe of patients undergoing monitoring for epilepsy surgery, a rare situation in which electrodes sit inside a conscious human brain for clinical reasons. They found number selective neurons. They also found something the monkey work could not have revealed: separate populations for dot arrays and for Arabic digits. Individual cells encoded one format or the other, not both [34]. A 2023 follow up found distinct neuronal representations for small and large numbers, with a coding boundary sitting close to four [35].
That boundary, appearing in human single cell recordings, is a striking convergence with a psychophysical elbow first noticed by a man throwing beans in 1871.
Not every imaging result supports a clean split. Piazza, Andrea Mechelli, Butterworth and Cathy Price used brain imaging to ask directly whether subitizing and counting recruit separate networks, and found a largely shared occipito parietal network rather than a tidy anatomical dissociation [36]. That result sits more comfortably with the continuum camp than with the dedicated mechanism camp.
Dehaene's broader framework, the triple code model, holds that the brain represents number in three formats: an analogue magnitude representation in the parietal lobe, a verbal representation tied to language areas, and a visual representation of written digits [39]. He and colleagues later specified three parietal circuits doing distinguishable jobs [38]. An early computational model built with Jean-Pierre Changeux had already shown that numerosity detectors of roughly this kind can emerge in a network without any explicit training on number [40]. The model has held up reasonably well, and the human single neuron findings on format specific coding fit it neatly.
What Babies Know, and What They Might Not
If number sense is built in, it should be visible before anyone teaches anything. Testing that on a preverbal infant requires ingenuity.
The standard tool is looking time. Babies look longer at things that surprise them. Show an infant the same display repeatedly and attention fades. Change something the infant cares about and attention returns. Whatever brings the gaze back is something the infant noticed.
Prentice Starkey and Robert Cooper used this logic in 1980 and reported that infants discriminated small numbers of dots [42]. Sue Ellen Antell and Daniel Keating extended it to newborns three years later [43]. Starkey, Spelke and Rochel Gelman then reported that infants match numbers across senses, looking longer at a display of two objects when they heard two drumbeats [44].
The most famous experiment came in 1992. Karen Wynn placed a single doll on a small stage, raised a screen, and visibly added a second doll behind it. Then the screen dropped. Sometimes two dolls stood there, as arithmetic requires. Sometimes only one, achieved by a hidden trapdoor. Five month old infants looked reliably longer at the impossible outcome, and the same held for subtraction [45]. The paper appeared in Nature under a title that made the claim unmistakable: addition and subtraction by human infants. Wynn later reported that infants enumerate actions as well as objects, extending the claim beyond things you can see all at once [46].
Véronique Izard, Coralie Sann, Spelke and Arlette Streri pushed the timeline to its limit in 2009. Newborns, roughly two days old, heard sequences of syllables and were then shown visual arrays. The babies looked longer at arrays whose number matched the number of sounds they had just heard [47]. Abstract number, the authors argued, before any experience of the world worth speaking of.
Fei Xu and Spelke separately showed that six month olds discriminate large sets, but only at a ratio of one to two [48]. Eight versus sixteen, yes. Eight versus twelve, no. Which is the approximate system showing its Weber signature at half a year old.
Now the objection, and it is a serious one.
Whenever the number of objects in a display changes, other things change with it. Two dolls have more total surface area than one doll. Three dots have more combined contour length than two dots. Denser arrays fill more of the visual field. An infant looking longer at a changed display might be responding to number, or might be responding to any of these correlated continuous quantities.
Melissa Clearfield and Kelly Mix designed a study to pull these apart. They habituated six to eight month olds to small sets, then tested them with displays that changed either number while holding contour length constant, or contour length while holding number constant. The infants dishabituated to the change in contour length. They did not reliably dishabituate to the change in number [49].
Mix, Janellen Huttenlocher and Susan Levine reviewed the whole infancy literature under a title that captures the problem exactly: multiple cues for quantification in infancy, is number one of them [50]. Feigenson, Susan Carey and Spelke then found that infants presented with a choice between tracking number and tracking continuous extent went with extent [51].
Wynn's dolls have also been reinterpreted. An infant who has built a mental file for each object and expects those files to persist behind a screen would show exactly the looking pattern Wynn reported, without doing any arithmetic. Object tracking, not addition.
The honest statement is this. Infants clearly notice something about quantity very early. Whether that something is number in any strict sense, or a bundle of correlated magnitudes that adults later disentangle, remains contested. The strong innateness claim, that newborns possess an abstract concept of number, is not established. It is a live hypothesis with real evidence on both sides.
The chronology above stops in 2024. That is not because the field went quiet. It reflects the limits of what could be verified for this article, and newer primary work almost certainly exists.
Two studies in mathematics education are worth noting because they are often cited loosely. A 2020 study of eighty children aged two to five found that those who could subitize to four, and who had also been taught counting, could reliably answer how many, while children who subitized only to three showed partial understanding [52]. A larger study following more than three thousand six hundred kindergarteners found that stronger subitizers had better arithmetic in first grade [53]. Both are correlational. Neither establishes that training subitizing produces the arithmetic gain.

One, Two, Many
In 2004, two papers appeared in the same issue of Science, both based on fieldwork in the Amazon, and together they changed how researchers think about the relationship between number words and number sense.
Peter Gordon of Columbia University had spent time with the Pirahã, a small group living along the Maici River in Brazil whose language has famously few number terms. Depending on how one analyses them, roughly one, two, and many. Gordon ran matching tasks. Lay out a row of objects, ask a participant to lay out the same amount. Performance was good for one, two and three. Beyond that it degraded sharply, and the errors grew in proportion to the size of the set, the signature of an approximate system rather than an exact one [54].
Pierre Pica, Cathy Lemer, Izard and Dehaene worked with the Mundurukú, another Amazonian group whose number words run out around five. Their results split cleanly. On approximate tasks, comparing two large sets or adding them roughly, the Mundurukú performed comparably to French controls, well beyond their naming range. On exact tasks, they failed for anything above four or five [55].

The conclusion drawn at the time was that number words are needed for exact arithmetic but not for approximate quantity. That is the version that circulated widely.
Four years later it was complicated. Michael Frank, Daniel Everett, Evelina Fedorenko and Edward Gibson returned to the Pirahã with a broader set of tasks. What they found was that Pirahã participants matched large sets perfectly when the target set stayed visible. Failures appeared only when the task required holding a quantity in memory or matching it out of sight. Their reframing has stuck: number words are a cognitive technology for tracking exact quantities across time and distance, not a prerequisite for perceiving quantity [56].
Caleb Everett and Keren Madora later ran further work among Pirahã speakers that broadly supported the technology framing [57]. And in a neat inversion, Frank and colleagues showed that English speaking adults, when their verbal resources are occupied by an interference task, start behaving in some respects like speakers of an anumeric language [58]. The exactness is in the words, and the words need a working verbal channel.
A parallel case comes from Australia. Butterworth, Robert Reeve and colleagues tested Aboriginal children whose languages have very restricted number vocabularies and found that they performed comparably to English speaking children on numerical tasks that did not require counting words [59].
Two cautions belong here. These are small field studies conducted under difficult conditions, with participant numbers in the dozens rather than the hundreds. And the wider debate about Pirahã grammar has been contentious in ways that go well beyond number. Treat the findings as important and suggestive rather than as settled.
What does this mean? Roughly that the counting system in your head has two layers. One you were born with, which handles small and approximate quantities and works without language. One you were given, made of words and symbols, which lets you carry an exact quantity across a room, across a week, or across a page.
Bees, Crows, and a Chimpanzee Named Ayumu
If number sense were a quirk of human cognition, it would be a curiosity. It is not a quirk.
Honeybees can be trained to match displays by number, generalising to novel stimuli they have never encountered [61]. More surprisingly, Scarlett Howard and colleagues at RMIT trained bees on a rule of choosing the lesser quantity and then presented them with an empty set. The bees treated empty as lower than one, placing zero at the bottom of a numerical continuum [60]. A concept that took human mathematics several thousand years to formalise appears to be available to an insect with roughly a million neurons.
Elizabeth Brannon and Herbert Terrace trained rhesus monkeys to touch arrays in ascending numerical order, then tested them on numbers from five to nine that they had never been trained on. The monkeys ordered them correctly [63]. That is not matching. That is generalising an ordinal rule.
Fish do it too. Christian Agrillo and colleagues at Padua showed that female mosquitofish placed in an unfamiliar tank prefer to join the larger of two shoals, and their discrimination follows a ratio dependent pattern [64]. In the wild, Karen McComb, Craig Packer and Anne Pusey played recordings of roaring intruders to lion prides in the Serengeti. The lionesses approached when they outnumbered the roars they heard and hung back when they did not [65]. Counting your enemies is not an academic exercise on the savannah.
And then there is Ayumu.
At the Primate Research Institute in Kyoto, Sana Inoue and Tetsuro Matsuzawa tested six chimpanzees on a task in which the numerals one through nine flash briefly on a touchscreen and are then masked by white squares. The subject must touch the squares in ascending numerical order from memory. Ayumu, a young male who had learned the task alongside his mother Ai, was tested against university students. At the shortest exposure, two hundred and ten milliseconds, Ayumu was both faster and more accurate than the human adults [62]. Six chimpanzees is a small sample, and the human participants were not trained to the same degree, both of which the authors acknowledged. The result has nonetheless held up as a demonstration that human numerical memory is not automatically superior.
Nieder has argued that the recurrence of numerical competence across insects, fish, birds and mammals reflects genuine adaptive value rather than incidental capacity [66]. Quantity matters for foraging, for territorial defence, for choosing a shoal, for deciding whether to fight. A brain that can tell more from less lives longer.
The convergence is the point. Crows and primates did not inherit number neurons from a common ancestor with number neurons. They arrived at similar solutions independently, in brains built on different plans. That is what an old and useful adaptation looks like.

Beyond the Eye
If subitizing were a property of the visual system, it should disappear when vision is removed. It does not.
Kevin Riggs and colleagues at London Metropolitan University built a device that raises small pins against a person's fingertips. Participants reported how many fingertips had been stimulated. Accuracy was ninety nine percent for one, ninety eight percent for two and ninety three percent for three, then dropped away sharply [67]. The same elbow, at a slightly lower number, in a sense that has nothing to do with the eye.
Myrthe Plaisier and colleagues in the Netherlands extended this to active touch, asking people to grasp a handful of objects and report the count. Subitizing appeared there too [68], and a later study found it across the fingers of a single hand [69].
Not everyone agrees. Alberto Gallace, Hong Tan and Charles Spence reanalysed the tactile literature and concluded that the case for genuine tactile subitizing is weaker than it appears, arguing that some of the effects can be explained without invoking a subitizing mechanism at all [70]. Their paper is titled as an unresolved controversy, and it remains one.

The most informative result came from Ludovic Ferrand, Riggs and Julie Castronovo, who tested congenitally blind adults on tactile enumeration. If subitizing depended on visual experience, people who had never seen anything should not show it. They did, performing comparably to sighted controls within the small number range [71]. Whatever the mechanism is, it does not need a lifetime of looking to develop.
Hearing shows the pattern too, with a lower ceiling. Valérie Camos and Barbara Tillmann presented sequences of tones and of visual flashes one at a time and found a discontinuity in enumeration performance at around two items [72]. Sequential presentation is harder than simultaneous, which makes sense if part of what the mind is doing is holding items open in parallel.
The pattern across the table is worth pausing on. The ceiling is not fixed at four. It is highest in simultaneous vision, lower in touch, lower still when items arrive one after another in time. That gradient suggests the limit is not a property of the eye but of some more central resource that different senses draw on to different degrees. It also connects to the way the mind briefly holds raw sensory input before anything is done with it, a topic covered separately in the work on sensory memory.
When the Count Breaks Down
Some people find basic numerical tasks disproportionately hard, in a way that does not track general intelligence, reading ability or educational opportunity. The condition is called developmental dyscalculia, and the research on it has become a testing ground for theories of number sense. What follows is about mechanisms and evidence, not about identifying or assessing anyone.
Butterworth's proposal, developed over two decades, is that dyscalculia stems from an impairment in a core capacity for representing quantity, sometimes called a defective number module [77]. The prediction is specific. If the core system for small exact quantities is impaired, then basic enumeration should be slow even in the subitizing range.
Karin Landerl, Anna Bevan and Butterworth tested thirty one eight and nine year olds identified as dyscalculic against controls matched for IQ, vocabulary and working memory. The dyscalculic group was slower on basic number tasks, including simple enumeration, despite the matching [73]. Patrick Schleifer and Landerl followed up with a study directly comparing subitizing and counting in typical and atypical development and found a narrower subitizing range in the atypical group [74].
Piazza and colleagues approached it from the approximate side, measuring number acuity across development and reporting that ten year olds with dyscalculia had Weber fractions comparable to typical five year olds [75]. Butterworth, Sashank Varma and Diana Laurillard summarised the case for a brain based account in Science [79], having earlier laid out the foundational capacities argument in more detail [78].

There is a serious competing account. Laurence Rousselle and Marie-Pascale Noël argued that the deficit is not in the quantity system itself but in accessing it from symbols. On their evidence, children with mathematics learning difficulties were impaired when comparing Arabic digits but not when comparing collections of dots, which points to a broken mapping rather than a broken core [76].
A third family of explanations does not invoke number at all, attributing difficulties to working memory, attention or inhibitory control. Given that attentional load demonstrably disrupts subitizing in typical adults, this is not a fringe position.
The honest summary. Diagnostic criteria vary between studies, samples are usually in the dozens, longitudinal work is scarce, and the population being studied is almost certainly not one thing. Different children may arrive at similar difficulties by different routes. The core deficit hypothesis is influential and supported, the access deficit hypothesis is a genuine rival, and the domain general accounts have not been excluded. Anyone presenting one of these as the established explanation is ahead of the evidence.
The Ceiling You Carry Everywhere
Step back and a curious coincidence appears. The number four keeps showing up in places that have nothing obviously to do with counting.
George Miller's 1956 paper put the span of immediate memory at seven plus or minus two, and became one of the most cited papers in psychology [80]. Miller's own point was subtler than the number suggests. What matters is not raw items but chunks, and a chunk can be as large as knowledge allows.
Nelson Cowan revisited the question in 2001 and argued that when you strip away rehearsal, grouping and long term memory support, the pure capacity of the focus of attention is closer to four [81]. Steven Luck and Edward Vogel had already put visual working memory at about four objects, and showed that those objects could each carry several features without extra cost [82]. Cowan returned to the theme in 2010 to ask why the number should be four at all [83].
So the subitizing ceiling, the multiple object tracking limit and the visual working memory capacity all cluster near the same value. Are they the same limit?
Possibly. The pointer account offers a tidy unification, with one pool of object files serving all three. But the identity has not been demonstrated, and there are dissociations that complicate it. Treat the convergence as a strong hint rather than a settled fact.
What is well established is what happens when the limit is respected rather than fought.
William Chase and Herbert Simon's chess work is the classic demonstration. Shown a board from a real game for five seconds, a master reconstructs it almost perfectly while a novice manages a handful of pieces. Shuffle the pieces into random positions and the master's advantage largely evaporates [84]. The master is not holding more items. The master is holding fewer, bigger ones. Fernand Gobet and colleagues later formalised how such chunks are built and retrieved [85].
The most striking illustration is a single participant. Anders Ericsson, Chase and Steve Faloon trained an undergraduate known as SF for more than two hundred hours on digit span. He began at seven digits, which is unremarkable. He finished near eighty. He did it by mapping digit groups onto running times, since he was a competitive runner, converting arbitrary strings into meaningful units [86]. His raw capacity never changed. Only the size of what fitted into it. This is the practical core of chunking in memory.
Numerosity research has produced its own version of this. Lorenzo Ciccione and Dehaene coined the term groupitizing for what happens when a display is arranged into small subgroups. Enumeration becomes faster and more accurate, and the improvement appears whether the grouping is done by spatial arrangement or by colour, which suggests people are exploiting arithmetic shortcuts rather than a purely perceptual trick [87].
Which explains a design convention nobody thinks about. Long numbers are broken into groups of two to four digits. Bank account numbers, telephone numbers, licence plates, the separators in a million. Playing cards and dominoes arrange their pips into subitizable clusters rather than rows. None of this is decoration. It converts a counting problem into a subitizing problem.
What does this mean for an adult trying to learn something difficult? Three things follow, and none of them involve training your subitizing range.
First, structure beats effort at the point of intake. A list of nine unrelated facts is nine items. The same nine facts organised into three groups of three is three items with internal structure, and the difference in what survives an hour later is not small.
Second, the limit is on simultaneous items, not on total knowledge. Expertise does not raise the ceiling. It raises what a single item can contain. This is why the same page of text is trivial to a specialist and impenetrable to a beginner, a phenomenon that cognitive load theory describes in detail.
Third, presentation format is not cosmetic. If the mind switches strategies based on how crowded a display is, then how information is laid out changes how much work is required to take it in. A table broken into small blocks is not merely prettier than a dense one. It is cheaper to read.

Notches in Bone
At some point, a mind that could hold four things decided to keep a record of more. When that happened, and what the earliest records mean, is where the evidence gets thin and the storytelling gets loose.
The safest artefact is the Lebombo bone, a baboon fibula recovered from Border Cave in the Lebombo Mountains on the border between South Africa and Eswatini. It carries twenty nine notches. Francesco d'Errico and colleagues reported twenty four radiocarbon determinations placing the organic material from that layer at roughly forty two thousand to forty four thousand years old [88]. The notches appear deliberate. Something was being marked.
What was being marked is another matter. Twenty nine is close to a lunar month, and that coincidence has generated a great deal of confident writing. It is a coincidence until proven otherwise. The bone is broken, so the original notch count is unknown.
The Ishango bone, from the Democratic Republic of Congo and roughly twenty thousand years old, is where interpretation has run furthest ahead of evidence. Its notches fall into three columns in grouped patterns, and over the decades these groupings have been read as evidence of prime numbers, of a duodecimal system, and of a six month lunar calendar. Specialists in the archaeology of notation are largely unpersuaded. The number of possible arithmetical patterns that can be extracted from three columns of tally marks is large, and finding one is not evidence that the maker intended it.
D'Errico's broader work on notched artefacts makes the methodological point. Deliberate marking is demonstrable through microscopy, showing that notches were made by different tools at different times, which indicates accumulation rather than decoration. What the accumulation counted is not recoverable from the object.

The defensible claim is narrow and still remarkable. By at least forty thousand years ago, humans were externalising quantity onto durable objects. A notch does not fade, does not need rehearsal, and does not care about a four item ceiling. It is the first move in a long escape from the limits of biological memory, and everything from tally sticks to written numerals to spreadsheets follows from it.
Whether any particular bone encodes mathematics is a separate question, and mostly an unanswered one.
Conclusion
Start with a die and end with a bone. In between sits a century and a half of experiments that agree on the phenomenon and disagree on the cause.
What can be said with confidence is a shorter list than most popular accounts suggest. Small sets are apprehended faster, more accurately and more confidently than large ones, and the transition is abrupt rather than gradual. The transition sits near four in adult vision, lower in touch, lower still in hearing, and moves with development. Something in the parietal and frontal cortex is tuned to numerosity, and equivalent tuning has evolved independently in birds. Quantity discrimination beyond the small range is ratio dependent in humans, monkeys, fish and infants. Number words are not required to perceive quantity, but they appear to be required to hold an exact quantity across time.
What cannot be said with confidence is almost as long. Whether subitizing is a dedicated mechanism or the low end of a single continuous system. Whether it operates before attention or depends on it. Whether infant looking times reflect number or correlated magnitude. Whether the four item ceiling in enumeration is the same ceiling as in working memory. Whether a core deficit in that system explains dyscalculia, or whether the difficulty lies in mapping symbols onto quantities, or somewhere else entirely.
That second list is not a failure. It is what a live field looks like from the inside.
There is something worth sitting with in all this. The ability that lets you glance at a die and know it is four is older than arithmetic, older than writing, older than the species. It does not improve with education. It does not respond to practice in any dramatic way. It sits underneath the numerical skills that were taught to you, unchanged, doing the same job it does in a crow.
Everything above that ceiling had to be invented. The words, the symbols, the notches, the notation, the place value systems that let a person hold a billion in mind without holding a billion of anything. That invention is what separates a mind that knows four from a mind that can calculate an orbit.
But the four came first. It came free. And on some level it is still the only number the brain gets for nothing.
Frequently Asked Questions
What is the difference between subitizing and counting?
Subitizing is the rapid apprehension of small quantities without any serial process, costing roughly 40 to 100 milliseconds per additional item. Counting is serial and costs around 250 to 350 milliseconds per item, requiring attention to move from object to object. The boundary between them sits near four items in adult vision.
Who first discovered subitizing?
The phenomenon was first measured systematically by economist William Stanley Jevons in 1871, using more than a thousand trials with beans thrown into a box. The term itself was coined in 1949 by Kaufman, Lord, Reese and Volkmann from the Latin subitus, meaning sudden. It was not coined by Piaget, despite that claim appearing widely.
Can animals subitize?
Numerical abilities appear across many species. Honeybees generalise by number and treat empty sets as lower than one. Rhesus monkeys order numerosities they were never trained on. Neurons tuned to number have been recorded in crows, whose forebrain evolved independently from the primate cortex, which points to convergent evolution rather than shared inheritance.
Is the subitizing limit exactly four?
No. It sits between three and four in adult vision and can reach five, but drops to roughly two or three in touch and around two for items presented one after another. It also expands during childhood. Spacing, density and attentional load all move it, so treating four as a fixed constant is misleading.
Does subitizing depend on attention?
This is genuinely unresolved. Studies loading attention with a secondary task show that small number accuracy degrades substantially, arguing against a preattentive process. Yet patients with parietal damage who cannot shift attention through a scene still enumerate one to three items accurately. Both findings have been replicated and neither camp has produced a decisive result.




