Introduction
Hold a sparkler in the dark and spin it fast enough and you see a circle. Not a moving dot. A circle. The light is only ever in one place at a time, yet your visual system stitches the whole loop together into something that looks continuous and solid. That ring is not out in the world. It is inside you, and it is the most ordinary demonstration of a memory system almost nobody thinks of as memory.
Sensory memory is the first holding stage between the world and your mind. It catches raw input from the senses, keeps it for somewhere between a tenth of a second and a few seconds depending on which sense you ask about, and then lets almost all of it go. What survives is whatever attention manages to grab before the trace dies. Everything else vanishes without ever reaching awareness.
Here is the part that makes it strange. For a brief window, you have access to far more detail than you can ever report. George Sperling proved this in 1960 with an experiment so clean it still gets taught in introductory psychology courses sixty-six years later [1]. He showed people a grid of letters for fifty milliseconds. They could name about four and a half of them. Then he changed one small thing about how he asked the question, and suddenly the same people could name nine. The letters had been there the whole time. The bottleneck was never the eye. It was getting the contents out.
This article follows that finding from a tachistoscope in a 1950s laboratory to single-neuron recordings in monkey visual cortex, and it does something the top-ranking pages on this topic mostly do not. It explains why the published numbers contradict each other. If you search for how long iconic memory lasts, you will find one hundred milliseconds, two hundred and fifty milliseconds, one second, and on at least one widely-read education site, five seconds. Those are not four opinions about one fact. They are four different measurements of four different things, and once you understand which paradigm produced which number, the confusion dissolves.

The Experiment That Started It All
George Sperling was a doctoral student at Harvard in the late 1950s working on a problem that sounds trivial and turned out not to be. How much can a person take in from a single brief glance?
The standard method was simple. Flash something. Ask what was seen. Count. And the answer had been stable for decades: about four or five items, regardless of how much you showed. Researchers called it the span of apprehension and treated it as a hard limit on perception itself.
Sperling did not believe it. His participants kept telling him the same thing after each trial. They had seen more than they could say. The letters were fading while they were still reporting the first ones.
So he built an experiment around that complaint. The results were published in 1960 as a full monograph in Psychological Monographs, seven experiments across twenty-nine pages, condensed from his dissertation [1].
The stimuli were cards bearing rows of letters and digits, viewed in a tachistoscope. Sometimes a single row of three to seven symbols. Sometimes a matrix, two rows of three, three rows of three, three rows of four. Standard exposure was fifty milliseconds.
The first thing worth noticing is what did not matter. Sperling varied exposure duration from fifteen milliseconds up to five hundred. Performance barely moved. Whether the array was on screen for a fifteenth of a second or half a second made almost no difference to how much people could report. Whatever was limiting them, it was not time on the retina.
Under whole report, where participants named everything they could, the average sat at roughly four and a half items. Exactly as the older literature said.
Then came the change that made the experiment famous. Instead of asking for everything, Sperling asked for one row. And critically, he only revealed which row after the display had already gone dark. A high, medium, or low tone sounded after offset, telling the participant which line to report. Since they could not know in advance which row would be tested, whatever they managed on the cued row should represent what they had available across the whole array.
The numbers jumped. On six-item cards, participants averaged about 2.8 correct from a three-item row, which scales to roughly 5.6 items available. On twelve-item cards, three rows of four, they averaged about 3.03 from a four-item row. Scale that up and you get around nine items available on a card where whole report had produced four and a half.
Twice as much information was sitting there. It simply could not survive the time it took to say it out loud.
Sperling then did the thing that turned a curiosity into a measurement. He varied when the tone arrived. Fifty milliseconds before the display went off. At the moment of offset. One hundred and fifty milliseconds after. Three hundred. Five hundred. One thousand.
The advantage decayed, and it decayed fast. It was largest when the cue came at or just before offset. It fell steeply across the first few hundred milliseconds. By a delay of one second the partial report advantage was gone entirely, and performance had collapsed back to the same four and a half items whole report always gave.
That decay curve is the shape of sensory memory. It is why the quarter second in this article's title is not a metaphor.
One detail deserves emphasis because most summaries skip it. The bottleneck Sperling identified was not storage. It was the transfer out of storage. Participants were not forgetting because the trace was weak. They were forgetting because reporting takes time, and the trace was dying while they reported. This is a distinction that matters enormously later, and it is closely related to why what you can recognise diverges so sharply from what you can recall.

The Marker That Erased What It Pointed At
A year after Sperling, two researchers at Bell Telephone Laboratories published a study that is cited far less often and arguably found something stranger.
Emanuel Averbach and A. S. Coriell were working on the same question with a different cue [2]. Instead of a tone indicating a row, they used a visual marker indicating a single letter position in a sixteen-letter array laid out as two rows of eight. Two marker types. A bar placed above the target position. Or a circle drawn around it.
With the bar marker, the story matched Sperling's. Information stayed accessible for roughly a quarter of a second, then faded. Storage time on the order of two hundred and fifty milliseconds. Capacity harder to pin down but clearly larger than what could be reported.
The circle marker did something else. It did not just point at the letter. Under certain timing conditions it destroyed it. Participants cued by a circle performed worse than participants cued by a bar, and worse than they would have with no cue at all.
Averbach and Coriell had discovered erasure. New visual information arriving at a location overwrites whatever was already stored there. The effect was local, tied to position, and it happens whether or not you want it to. This is backward masking, and it is one of the most important facts about sensory memory for anyone who cares about how people actually take in visual material.
Worth noting for honesty: a 1964 partial replication by Mayzner and colleagues failed to reproduce the delay-dependent recognition drop [55]. These paradigms are sensitive to procedural detail in ways the textbook summaries rarely admit.

The Man Who Named It
Neither Sperling nor Averbach used the word iconic. That came from Ulric Neisser, whose 1967 book Cognitive Psychology did more than almost any single work to define the field it named [3].
Neisser gave the visual store a name, iconic memory, and gave the auditory equivalent one too, echoic memory. He framed both as brief, modality-specific, precategorical buffers sitting before pattern recognition. Raw, unprocessed, not yet sorted into meaning.
He did not run the founding experiments. He synthesised them, named them, and in doing so gave a generation of researchers a shared vocabulary. That is not a small contribution. But it is worth being clear about what happened, because a lot of popular summaries imply Neisser discovered these stores. He organised them.
The precategorical claim, that these buffers hold raw sensation with no meaning attached, has aged less well than the names. More on that later.
The Three-Eared Man
If vision has a brief store, hearing should too. Testing it turned out to be harder than it sounds.
The problem is that Sperling's design uses a tone to cue a visual display. You cannot cue an auditory display with another sound without the cue itself interfering with what you are trying to measure. Christopher Darwin, Michael Turvey and Robert Crowder solved this in 1972 with a design that has been called the three-eared man [4].
They built three spatially distinct auditory channels using headphones. One appearing to come from the left. One from the right. One centred, presented to both ears at once, which the brain localises as coming from inside the head. Each channel carried a short list of items. Three channels, three items each, giving a nine-item auditory analogue of Sperling's twelve-letter array.
The cue was visual. A light indicating which spatial stream to report, delivered at various delays after the sounds ended.
The first experiment found exactly what the visual work predicted. Partial report beat whole report, and the advantage held when the cue came less than about four seconds after offset. Capacity settled at roughly four to five items.
A second experiment tried cueing by category instead of location, letters versus digits. Here partial report gave no advantage over whole report when participants also had to report location. When location was not required, the advantage returned but only for delays under about two seconds.
So the auditory store lasts longer than the visual one and holds less. Two to four seconds versus a quarter of a second. Four or five items versus nine or more. That asymmetry makes functional sense. Vision can re-sample the world by simply looking again. Sound arrives once and is gone, so the system that catches it has to hold on longer.
Donald Massaro had reached a compatible conclusion two years earlier working on what he called preperceptual auditory images, arguing that a brief auditory representation must persist long enough for recognition to operate on it [42].

Two Thousand Years of Chasing an Afterimage
None of this started in 1960. People have been noticing that images linger for as long as anyone has written about perception. Aristotle described afterimages persisting after the stimulus was gone, including in dreams. In the eighteenth century Segner spun a glowing ember on a wheel and used the point at which the trail closed into a complete circle to estimate how long visual persistence lasts. Wilhelm Baxt used masking in 1871 to limit how much could be extracted from a brief presentation. Coltheart's 1980 review remains the best single source for this pre-history [5].
Notice what happened across those two millennia. The question shifted from does anything linger, which was settled early, to what exactly is lingering, which is still being argued about.
Why Every Source Gives You a Different Number
Search for the duration of iconic memory and you will get a mess. One hundred to two hundred milliseconds. Two hundred and fifty milliseconds. Half a second. One second. Up to five seconds. Echoic memory is worse: two to three seconds, two to four seconds, up to ten seconds, and in at least one place twenty.
This looks like a field that cannot get its story straight. It is not. Almost every one of those numbers is defensible. They are measurements of different constructs produced by different paradigms, and the sources quoting them almost never say which.
Here is the reconciliation, which as far as I can tell does not exist anywhere else on the first page of results for this topic.
| Register | Reported duration | Capacity | Paradigm that produced the number | Source |
|---|---|---|---|---|
| Iconic, informational persistence | ~250 ms, advantage gone by ~1 s | ~9 items available vs ~4.5 reportable | Partial report with post-offset cue, 50 ms exposure | Sperling 1960 |
| Iconic, visible persistence | ~130 ms maximum, often much less | not separately estimated | Temporal integration and synchrony judgements. Inversely related to stimulus duration and luminance | Coltheart 1980; Di Lollo 1980 |
| Iconic, buffer with erasure | ~250 ms | high, above report limit | Bar and circle markers with backward masking | Averbach and Coriell 1961 |
| Iconic, item retention vs precision | items lost within ~1 s | precision of survivors stable | Continuous report with mixture modelling | Pratte 2018 |
| Echoic, short store | up to ~300 ms | extends apparent stimulus duration | Recognition and stimulus-extension tasks | Cowan 1984 |
| Echoic, long store | 2 to 4 s behaviourally, ~10 s neurally | ~4 to 5 items | Auditory partial report; mismatch negativity across interstimulus intervals | Darwin et al. 1972; Sams et al. 1993 |
| Haptic | ~0.8 s, up to ~10 s with rehearsal suppressed | ~4 to 5 items | Air-jet stimulation, partial vs whole report; interference tasks | Bliss et al. 1966; Gilson and Baddeley 1969 |
| Olfactory | not reliably established | very low | Paired-succession similarity and delayed matching | White 1998 |
| Gustatory | no dedicated evidence base | unknown | none comparable exists | see text |
Read down the iconic rows and the puzzle resolves. Visible persistence is the phenomenal experience of the stimulus still being there. Informational persistence is the availability of readable content, whether or not it feels visible. They are not the same thing, they behave differently, and one of them behaves in a way that surprises almost everyone.

Three Kinds of Persistence, Not One
In 1980 Max Coltheart published a review in Perception and Psychophysics that reorganised the entire field [5]. His argument was that researchers had been treating one word, persistence, as if it named one thing, when it actually named three.
Neural persistence is continued afteractivity in the visual pathway, from retina onward. Physiology.
Visible persistence is the phenomenal experience of still seeing something that is no longer there. And here is the counterintuitive part. Visible persistence is inversely related to how long and how brightly the stimulus was presented. Longer stimulus, shorter persistence. Brighter stimulus, shorter persistence. Vincent Di Lollo demonstrated this with temporal integration tasks in the same year and found that persistence drops toward zero once stimulus duration passes roughly one hundred to one hundred and fifty milliseconds [6].
Informational persistence is the availability of content that partial report can extract. It is roughly independent of exposure duration, which is exactly what Sperling found when varying exposure from fifteen to five hundred milliseconds changed almost nothing.
Coltheart's conclusion was blunt. Iconic memory, meaning informational persistence, is not the same as visible persistence, and methods that ask people what they still see do not validly measure it. A generation of studies had been measuring the wrong construct and calling it the right one.
David Irwin and James Yeomans refined the picture in 1986 [7]. Testing three-by-three arrays across exposures from fifty to five hundred milliseconds, they found evidence for a visual memory that begins at stimulus offset and lasts roughly one hundred and fifty to three hundred milliseconds, largely independent of how long the display had been up. Different clock, different mechanism.
Robert Efron had already established something related in 1970 from a completely different angle. He found that both auditory and visual perceptions have a minimum duration, produced by stimuli of about one hundred and twenty to one hundred and thirty milliseconds or less, with perceptions from longer stimuli graded continuously [53]. The perceptual system has a floor.
Coltheart's three-way split also settles one of the most repeated claims about vision, and it is worth settling carefully, because the sparkler at the top of this article is the one case where the claim is actually true.
The myth goes like this. Film works because visible persistence holds each frame in your eye long enough to bridge the gap to the next one. The gaps fill in, and motion appears. Nearly every popular explanation of cinema says some version of that. It is wrong.
Visible persistence does have a job in a cinema, and the job is not motion. It smooths flicker. Successive frames and the dark intervals between them stop registering as flashing, which is why a projected image looks steady rather than strobing. That is flicker fusion, and persistence genuinely contributes to it.
Perceived motion runs on different machinery. Apparent motion is produced by dedicated motion-detecting processes that compare position across time, not by a lingering image plugging a hole. Oliver Braddick identified a short-range process operating over small spatial and temporal separations in 1974 and distinguished it from a longer-range process working across larger displacements [57]. Stuart Anstis reviewed the mechanisms at the end of that decade [58]. The dichotomy has since been contested in its own right. Patrick Cavanagh and George Mather argued in 1989 that the short-range and long-range differences follow from the stimuli each paradigm uses rather than from two qualitatively different systems [59].
That argument is still running. What is not in dispute is the piece the popular account has backwards. If persistence were doing the work, a longer or brighter frame would give you more of it. Di Lollo showed the opposite. Persistence shrinks as duration and luminance rise, so brighter projection would make cinema worse rather than better.
The sparkler is genuine visible persistence. The trail of a single moving point really does linger, and the ring really is inside you. The film is not the same thing. Two everyday demonstrations that feel identical, and only one of them is what people think it is.

The Store That May Not Be a Store
Every textbook diagram shows three boxes. Sensory memory feeds short-term memory feeds long-term memory. Clean, teachable, and probably wrong as a description of how the brain is actually organised.
Nelson Cowan has been the most persistent critic. In 1984 he published a review arguing that auditory sensory memory is not one store but two phases [8]. A short auditory store that extends the apparent duration of a stimulus up to about three hundred milliseconds, used in recognition. And a long auditory store retaining information for at least several seconds, with acoustic detail fading over something like ten to fifteen.
That single paper resolves most of the echoic contradiction in the table above. Two to four seconds and ten seconds are both right. They refer to different phases.
Four years later Cowan went further and questioned the boxes themselves [9]. In his embedded-processes account, working memory is not a separate container. It is the currently activated portion of long-term memory, with the focus of attention nested inside that activated set. Sensory memory, in this view, is not a waiting room outside the system. The longer phase of it just is activation of features already represented in long-term memory.
That reframing has consequences for the precategorical claim. If the store is partly activated long-term representations, it cannot be entirely free of meaning. And the empirical support for precategoriality was always shakier than the textbooks suggest. Category-based partial report cueing, letters versus digits, works poorly, which Darwin and colleagues found in their own second experiment [4].
A sharper attack came in 2016. Haluk Öğmen and Michael Herzog argued that the classic retinotopic iconic store cannot hold anything useful under normal viewing conditions [10]. The reasoning is hard to dismiss. Sperling's participants stared at a fixed point while a display flashed. In ordinary life your eyes move three or four times a second and objects move independently. A store organised by retinal position would be scrambled by the first saccade. They proposed an additional non-retinotopic component organised by motion grouping rather than retinal location. Tripathy and Öğmen extended this in 2018, arguing sensory memory is allocated to the current event segment rather than to a fixed retinal map [54].
There is also evidence for something in between. Ilja Sligte and colleagues have described a fragile visual short-term memory, an intermediate store with several seconds of duration and much higher capacity than working memory, sitting between the icon and working memory proper [11], [12]. Three boxes was never going to survive contact with the data.
Where does that leave the concept? Useful, but not literal. Sensory memory names a real set of phenomena. It probably does not name a distinct anatomical container. The better description is modality-specific neural and perceptual persistence, plus activation of existing representations, with attention doing the selecting. Which is less tidy than three boxes and closer to true.
Iconic Memories Die a Sudden Death
For decades the assumed picture of iconic decay was a fade. The icon dims. Detail blurs. Eventually there is nothing left. It matches the phenomenology and it matches the sparkler.
Michael Pratte tested it directly in 2018 and found something different [13].
The method is elegant. Participants saw ten coloured squares arranged in a circle for two hundred milliseconds. After a variable delay, from about thirty-three milliseconds up to a full second, an arrow pointed at one location and they reported the colour that had been there by selecting from a continuous colour wheel rather than a fixed set of options.
The continuous response is what makes the study work. With a multiple choice answer you only learn whether the participant was right. With a continuous colour wheel you learn how wrong they were, and mixture modelling can then separate two things that had always been confounded: the probability that an item was retained at all, and the precision of the items that survived.
The dissociation was clean. As delay increased, items were lost completely. But the precision of whatever remained was only marginally affected by the passage of time. Iconic memory does not blur. It loses whole items, all or nothing, while the survivors stay sharp.
The title Pratte gave the paper says it well. Iconic memories die a sudden death.
Two honest caveats, because this is exactly the kind of finding that gets over-reported.
First, a corrigendum was published later that year in the same journal [14]. It corrected a technical detail in the reporting. It did not change the conclusions.
Second, and more important, I could not find a large independent replication. The result fits well with other work from the discrete-capacity tradition, including Pratte's own 2017 study on stimulus-specific variation in precision [15] and Rouder and colleagues' assessment of fixed-capacity models [16]. But fitting an existing theoretical camp is not the same as being independently replicated. Treat it as an influential single-laboratory finding, not settled fact. The slot-versus-resource debate about visual memory capacity is very much still open.

What the Electrodes Show
Everything so far comes from asking people what they saw or heard. There is another way in, and it does not require the participant to report anything at all.
In 1978 Risto Näätänen, A. W. K. Gaillard and S. Mäntysalo published a reinterpretation of an evoked potential effect that turned out to be one of the most productive findings in cognitive neuroscience [17]. Play a repeating standard tone. Occasionally slip in a deviant. Roughly one hundred to two hundred and fifty milliseconds after the deviant, a negative deflection appears in the difference wave, largest at fronto-central electrodes.
They called it the mismatch negativity, or MMN. The remarkable part is that it appears whether or not the listener is paying attention. Participants can be reading a book or watching a silent film. The brain flags the deviation anyway.
Näätänen's interpretation was that MMN reflects a comparison against a short-term auditory memory trace of the standard. If no trace exists, no comparison happens, and no MMN appears. A single deviant with no preceding standards produces nothing.
That interpretation makes a testable prediction, and it holds. As the interval between stimuli lengthens, MMN amplitude falls and its latency increases, which is what you would expect if the trace against which comparison is made is decaying. Böttcher-Gandor and Ullsperger mapped this across varying interstimulus intervals in 1992 [20]. Sams and colleagues used magnetoencephalography in 1993 and put the persistence of the trace at about ten seconds [19], considerably longer than the two to four seconds behavioural partial report gives. Both are measuring something real. They are not measuring the same thing.
The 2007 review by Näätänen, Paavilainen, Rinne and Alho remains the standard reference for the paradigm and its parameters [18]. A systematic review by Bartha-Doering and colleagues in 2015 pulled together thirty-seven studies specifically treating MMN as an index of auditory sensory memory [36].
Now the complication, which almost no popular source mentions.
There is a competing account. When a standard repeats, the neurons responding to it adapt. Their response weakens. A deviant, being different, hits a less-adapted population and produces a stronger response. On this reading MMN is not a memory comparison at all. It is differential adaptation.
Thomas Jacobsen and Erich Schröger designed the control that addresses this in 2001 [21]. The many-standards or equiprobable paradigm presents a set of tones each occurring equally often, so that no single tone is repeated enough to build up adaptation. If MMN survives against that control, it cannot be pure adaptation. It largely does, which supports a genuine memory-based comparison component.
Marta Garrido, James Kilner, Klaas Stephan and Karl Friston went further in 2009, reconciling the adaptation and memory-trace accounts under predictive coding [22]. In that framework MMN is a hierarchical prediction error signal. The brain builds a model of the regularity, and the deflection reflects the mismatch between prediction and input. Näätänen himself had described the underlying process as a kind of primitive intelligence in auditory cortex [23]. A 2025 review by Fisher and Todd surveys where the field stands after nearly fifty years [56].
So MMN is a window onto auditory sensory memory. It is not a clean readout of it. Anyone telling you the MMN simply measures how long echoic memory lasts is skipping an argument that has been running for two decades.

The Gate That Is Not a Memory
There is a second electrophysiological measure that gets confused with the first constantly, and the confusion is worth clearing up because the two measure genuinely different processes.
In 1982 Lawrence Adler and colleagues published a study using a paired-click paradigm [24]. Two identical clicks separated by about half a second. In a healthy listener, the cortical response to the second click is dramatically smaller than to the first. The brain has recognised the stimulus as redundant and suppressed its own response.
The numbers from that paper are stark. At the half-second interval, healthy controls showed over a ninety percent mean reduction in response to the second click. Participants with schizophrenia showed less than fifteen percent. At two-second intervals, responses in controls were still thirty to fifty percent reduced.
This is the P50 suppression effect, and it indexes sensory gating. The automatic inhibition of responses to repeated, uninformative input. It is tied to cholinergic and nicotinic receptor systems.
Gating is not memory. A memory trace holds content for comparison. A gate decides how much of the incoming signal gets through in the first place. Both are automatic, both are pre-attentive, both are measured with scalp electrodes, and both are impaired in schizophrenia. That is where the conflation comes from. But if you want to understand sensory memory specifically, MMN is the relevant measure and P50 is a different question. Meta-analytic work by Bramon and colleagues confirms the P50 deficit in schizophrenia is reliable [25], with effect sizes in patient samples typically in the moderate to large range rather than the implausibly enormous figures that circulate on aggregator sites.
Watching the Icon in a Living Brain
Scalp electrodes average across millions of neurons. To see what a sensory trace looks like at the cellular level you need to record from inside, which means animal work.
In 2021 Rob Teeuwen, Catherine Wacongne, Ulf Schnabel, Matthew Self and Pieter Roelfsema published the first direct test linking iconic memory to activity in primary visual cortex [26]. They measured iconic memory behaviourally in macaques, quantifying what they called its worth: how many extra milliseconds of viewing time the lingering icon is effectively equivalent to, computed by asking how much longer a masked stimulus would need to be shown to match performance on an unmasked one.
Then they looked at V1. Neural activity persisted after the stimulus disappeared, and that persistent activity predicted whether the animal would report correctly. The time course of the neural persistence matched the time course of the behavioural worth and decay. A post-offset attention cue boosted responses only while the icon was still alive.
That is about as direct a neural correlate as this field has produced. It is also two monkeys, which is normal for single-unit physiology and still worth stating plainly.
Higher up the visual hierarchy, Christian Keysers, Dengke Xiao, Peter Földiák and David Perrett recorded from the superior temporal sulcus using rapid serial visual presentation [27]. When gaps were inserted between successive images, neurons in STS continued processing a stimulus as though it were still present for gaps of up to about ninety-three milliseconds. Their earlier work on the speed of sight established the timing framework for these experiments [28].
Two honest limits are worth stating. The claim that sensory memories are stored in the thalamus turns up in several popular health and psychology pages, and the evidence for it is thin. MMN generators localise to auditory cortex together with inferior and medial frontal regions across surface EEG, intracranial recording and imaging. The thalamus is involved in relaying sensory input, which is not the same as being the store.
And more broadly: the neuroscience shows persistent, modality-specific activity that correlates with behavioural measures of sensory memory. It does not show a dedicated anatomical structure whose job is holding the icon. That distinction gets flattened constantly in secondary sources.

The Register Changes Across a Lifetime
Sensory memory is not a fixed constant. It develops, and it declines.
Elisabeth Glass, Steffi Sachse and Waldemar von Suchodoletz used MMN to map auditory sensory memory duration across early childhood in 2008 [29]. Memory traces for tone characteristics lasted one to two seconds in two and three year olds, more than two seconds in four year olds, and three to five seconds in six year olds. The register grows.
This has a practical edge. Short auditory sensory memory duration in young children is associated with delayed language development. If the trace fades before the next syllable arrives, building words out of a sound stream becomes much harder. Näätänen and colleagues surveyed the clinical range of these findings in 2012 [35].
At the other end of life, the trace shortens again. Eero Pekkonen's review of MMN in ageing and in Alzheimer's and Parkinson's diseases documents faster decay at longer interstimulus intervals with age, with a markedly shortened trace in Alzheimer's [30].
The visual side shows a parallel pattern. Zhong-Lin Lu and colleagues found faster decay of iconic memory in observers with mild cognitive impairment [31]. The icon dies sooner.
These are group-level differences with real overlap between individuals. None of them functions as a diagnostic test on its own.
When the Trace Breaks Down
The best-established clinical finding in this literature concerns schizophrenia, and the effect size is large enough to be worth stating precisely.
Daniel Umbricht and Sanya Krljes published a meta-analysis of thirty-two studies in 2005 [32]. The mean effect size for MMN reduction in schizophrenia was 0.99, with a ninety-five percent confidence interval from 0.79 to 1.29. In a field where much smaller effects get celebrated, that is unusual.
The pattern across illness stage matters. Molly Erickson, Abigail Ruffle and James Gold analysed the progression in 2016 [33]. Reductions are small in people at clinical high risk, moderate at first episode, and large in chronic illness. Duration-deviant MMN is the most consistently impaired variant.
Dean Salisbury and colleagues added an important refinement in 2020 [34]. More complex, pattern-based deviants show reductions earlier in the illness than simple tone deviants do, with a substantial effect already present at first episode. Related work has linked MMN reduction in chronic illness to grey matter volume in Heschl's gyrus, the primary auditory region.
On specificity: effects are smaller in bipolar disorder and in first-episode psychosis generally than in chronic schizophrenia. MMN is not a schizophrenia detector. It is a sensitive index of auditory processing integrity that happens to be badly affected in that condition.
Can Sensory Memory Be Trained?
This question deserves its own section because the answer circulating online is wrong, and the correct answer requires a distinction that almost nobody makes.
There are three different claims hiding under the phrase improving your sensory memory.
One. Extending the raw duration or capacity of the register itself. There is no good evidence for this. No study shows that training lengthens the fundamental iconic window of roughly a quarter second to one second, or the echoic window of two to four seconds, in healthy adults. The duration appears to be a property of how the sensory systems are built.
Two. Improving discrimination and attentional selection downstream. This is genuinely trainable. Claudia Lappe and colleagues showed that short-term musical training enhances MMN amplitude to trained deviants, with sensorimotor training producing larger effects than listening alone [37]. Bonnie Lau and colleagues found that brief perceptual learning produces sustained changes in cortical evoked responses [38]. What improves is how well the system encodes, represents and selects, not how long the buffer holds.
Three. Far transfer to general cognition. Not supported. Perceptual learning is largely specific to the trained stimuli. Getting better at discriminating trained pitch contours does not make you better at remembering names.
So the accurate statement is this. You cannot make the register last longer. You can get substantially better at pulling material out of it before it dies, because attention is the gate between the sensory register and everything downstream, and attention is trainable. That is a real and useful conclusion. It is just not the one usually advertised.
The Photographic Memory Confusion
The other thing readers merge with sensory memory is photographic memory. The two share almost nothing beyond the vague sense of an image outlasting the thing that caused it.
Iconic memory is universal. Everybody has it. It runs for roughly a quarter of a second, it holds material that has not been categorised yet, and nobody experiences it as a picture they can sit and inspect. It is gone before you could decide to look.
Eidetic imagery is a different animal. Ralph Norman Haber and Ruth Haber tested nearly every child in an elementary school in New Haven in 1964, using strict criteria to separate genuine eidetic images from afterimages and from ordinary recall [60]. Twelve children, about eight percent, met every criterion. Their images lasted as long as four minutes. Not milliseconds. Minutes.
Then comes the result that should have ended the photographic memory story on the spot. Once the imagery faded, those children remembered the picture no better than the children with no eidetic imagery at all. The image was real. It simply was not doing anything for retention.
Cynthia Gray and Kent Gummerman went back through the methods and the data a decade later and found the evidence far shakier than the popular version suggested [61]. Haber returned to the topic himself in 1979 with a title that tells you how those twenty years had gone, asking where the ghost was [62]. Eidetic imagery is rare, largely confined to young children, and mostly gone by adolescence, and even among those who show it only a small fraction produce anything close to an accurate reproduction of the original.
So: sub-second and universal on one side, minutes-long and rare and developmental and no help to memory on the other. Whether anything deserving the name photographic memory exists in adults at all is a separate question with its own evidence base.
The Registers Nobody Can Measure
Almost every article on this topic gives you five sensory registers with confident numbers attached. Iconic, echoic, haptic, olfactory, gustatory. Sight, sound, touch, smell, taste. Neat and symmetrical.
Two of those five are not supported by the evidence base the way the other three are, and I think saying so plainly is more useful than pretending otherwise.
Haptic memory is real and reasonably well studied, but the numbers vary more than they should. James Bliss and colleagues at Stanford Research Institute in 1966 used air jets delivered to the finger joints and ran the partial-report logic on touch, finding a store lasting under a second [39]. Elizabeth Gilson and Alan Baddeley in 1969 found tactile information could persist far longer, up to ten seconds, when verbal rehearsal was suppressed [40]. That gap between under a second and ten seconds is not noise. It reflects whether the participant is allowed to recode the sensation into words. Which raises the question of whether the longer figure is measuring a sensory register at all.
Olfactory memory is where confidence outruns data. Theresa White's 1998 review in Chemical Senses found only preliminary support for a short-term odour store distinct from other memory systems [41]. Most olfactory memory research concerns long-term associative memory, which is genuinely striking, rather than a Sperling-style precategorical buffer.
Gustatory sensory memory has essentially no founding literature. There is no taste equivalent of Sperling 1960 or Darwin 1972. The row exists in tables because the symmetry looks right.
If you see a source giving you a crisp figure for how many milliseconds taste sensory memory lasts, that number came from somewhere other than a study.
What This Actually Means for Studying
Everything above is descriptive science. But the sensory register is the first checkpoint every piece of material you study has to pass, and its properties place real constraints on how study material should be arranged. Almost no popular article on this topic makes the connection, so here it is.
Background sound is not neutral, and which sound matters more than how loud. Pierre Salamé and Alan Baddeley showed in 1982 that unattended speech disrupts serial recall of visually presented items [44], building on earlier work by Colle and Welsh [43]. Dylan Jones and colleagues then found the variable that actually governs the effect, and it is not meaning or volume. It is acoustic change [45]. A repeating steady sound barely disrupts anything. A changing sequence disrupts a lot. Pure tones disrupt about as much as speech [46], and disruption grows with the variety of sounds in the set [47].
There is an important nuance here that changes the practical advice. The disruption appears in tasks requiring serial order and not in tasks like missing-item recall [48], which means the effect operates on order processing in working memory rather than on the raw sensory trace. Hughes and colleagues have argued for a duplex account with both interference-by-process and attentional capture contributing [49]. Practically: acoustically varied background audio costs you most when the material has order that matters, such as sequences, steps, spellings and formulas. This is one of the better-evidenced reasons that the environment you study in changes what you retain.
Fast visual streams erase themselves. Averbach and Coriell showed that new information arriving at a location wipes out what was stored there. Keysers and colleagues found temporal cortex treats successive images as continuous only up to about ninety-three millisecond gaps. Scroll a feed fast enough and each item masks the one before it before attention has selected anything from it. The material never fails to be encoded because it was boring. It fails because it was overwritten. The same logic explains part of why interruptions are so costly to learning.
Pace governs how much survives. Geoffrey Loftus, Janine Duncan and Paul Gehrig demonstrated in 1992 that the perceptual information extracted from a brief presentation is governed by total processing time, meaning stimulus onset to mask onset, rather than raw exposure duration [50]. Translated to a lecture or a slide deck: what matters is not how long a slide is technically visible but how long it remains available before the next one displaces it. Slower transitions on dense slides are not indulgence. They are the difference between transfer and erasure.
Visual clutter costs you throughput. The register holds a lot but attention can only extract a few items before decay. More competing elements means fewer relevant ones make it out. Which is also why grouping material into meaningful units pays off so reliably.
Splitting across senses buys capacity. Seyed Mousavi, Renae Low and John Sweller showed in 1995 that presenting some information through the ear and some through the eye reduces load and improves learning compared with routing everything visually [51]. Roxana Moreno and Richard Mayer replicated and extended the modality principle in 1999 [52]. The registers are modality specific, so they do not compete for the same capacity. The corollary matters just as much: reading identical text aloud while it is on screen makes things worse, not better, because both versions land in the same channel. This sits directly on top of the way load accumulates across channels, and it has clear implications for anyone whose job depends on holding several pieces of information at once under time pressure.

Two Numbers on the First Page of Google That Are Wrong
Search this topic and you will land on two claims that the primary literature does not support. Both come from sites with real traffic.
Claim one: iconic memory lasts up to five seconds.
It does not. Iconic informational persistence runs about two hundred and fifty milliseconds, with the partial-report advantage gone by roughly one second [1]. Visible persistence is shorter still, topping out near one hundred and thirty milliseconds and shrinking as stimulus duration and brightness increase [5], [6].
Five seconds is roughly the figure for echoic memory, and even there the standard behavioural estimate is two to four seconds [4]. The five-second number appears to be echoic memory misfiled under the visual heading. The same page also states that sensory memory lasts less than a minute, which is technically true and about as useful as saying a sprint takes less than a day, and repeatedly writes hepatic where haptic is meant. Hepatic pertains to the liver.
Claim two: you can train and strengthen your sensory memory.
Not as stated. The evidence supports improved perceptual discrimination and attentional selection after training [37], [38]. It does not support lengthening the register itself, and it does not support transfer to unrelated tasks. The honest version of the claim is narrower and still worth knowing: you can get better at extracting information before it decays.
Neither of these is malicious. Both come from compressing a nuanced literature into a listicle. But when a student cites the five-second figure in an exam answer, they are citing a page, not a study.

Conclusion
Start with the sparkler. That ring of light is your visual system holding onto something that is no longer there, and it is the closest most people ever come to noticing sensory memory in action.
The science underneath it is less tidy than the textbook diagram suggests. Sperling proved that far more information is briefly available than can ever be reported, and that the bottleneck sits in getting it out rather than getting it in. Averbach and Coriell showed the trace can be erased by whatever arrives next in the same place. Neisser named the visual and auditory versions. Darwin, Turvey and Crowder built the auditory analogue and found a store that lasts longer and holds less. Coltheart then showed that persistence was three different phenomena wearing one word, which is the single most useful idea in this entire literature and the reason the published numbers appear to contradict each other when they mostly do not.
Since then the picture has shifted further from the three-box model. Cowan reframed the longer phase of sensory storage as activation of long-term representations rather than a separate container. Öğmen and Herzog pointed out that a retinally-organised store cannot survive normal eye movements. Sligte and colleagues found something with several seconds of duration sitting in between the icon and working memory. Pratte found that the icon does not fade so much as drop items outright, all or nothing, while whatever survives stays sharp.
The electrophysiology adds a layer the popular accounts almost entirely skip. Mismatch negativity gives a measure of the auditory trace that does not require the listener to report anything, and it changes across development, across ageing, and in schizophrenia with an effect size around 0.99 that most of psychology would envy. It is also not a clean readout, because adaptation and prediction error contribute to the same signal, and that argument is not finished.
What should you take from this if you are not a memory researcher? Three things. The register is brief and largely fixed, so you cannot train it longer. Attention is the gate, and the gate is trainable. And the conditions you study under, the acoustic variability in the background, the pace at which material is replaced, the visual density of what you are looking at, all operate on this first checkpoint before anything else in learning gets a chance.
Everything you see is still there for a moment after it is gone. Almost none of it survives. What determines which fragments make it through is not how hard you try to remember. It is what you were pointed at, and how fast the next thing arrived.
Frequently Asked Questions
How long does sensory memory last?
It depends on the sense and on how the question is measured. Visual iconic memory lasts roughly a quarter of a second, with the partial report advantage gone by about one second. Auditory echoic memory lasts two to four seconds behaviourally. Haptic estimates range from under a second to several seconds.
What is the difference between iconic and echoic memory?
Iconic memory is the visual sensory register. It holds a large amount of detail for around 250 milliseconds. Echoic memory is the auditory register. It holds fewer items, roughly four or five, but keeps them noticeably longer, typically two to four seconds. Sound arrives once, so the system holds it longer.
What did Sperling's experiment actually prove?
It proved that far more information is briefly available than a person can report. Under whole report participants named about 4.5 letters. Cued after the display vanished, the same participants performed at a level implying about nine letters were available. The limit was report speed, not perception.
Can sensory memory be improved with training?
Not its raw duration. No evidence shows training lengthens the iconic or echoic window in healthy adults. Training does improve perceptual discrimination and attentional selection, which means more information gets extracted before the trace decays. Claims about broad transfer to unrelated cognitive tasks are not supported.
Is sensory memory a real separate store in the brain?
Increasingly this is debated. The classic model treats it as a distinct stage before short-term memory. Cowan argued the longer phase is activated long-term memory rather than a separate container, and later work questioned whether a retinally organised visual store could function during normal eye movement.




