Introduction
On the morning of September 12, 2001, a graduate student walked into classrooms at Duke University and handed out a questionnaire. Fifty-four students wrote down where they had been when they first heard about the attacks the previous morning. Who told them. What they were doing. What they felt. Then, almost as an afterthought, they were asked to describe an ordinary event from the days before, something forgettable, a party or a study session or a walk across campus [1].
That second question is the reason this study matters. It gave the field something it had lacked for twenty-four years: a control.
Flashbulb memories are the recollections people carry of the moment they learned about a shocking public event. The term was coined in 1977 by two Harvard psychologists who thought they had found something extraordinary, a class of memory so vivid and so durable that it seemed to work by different rules [2]. Nearly fifty years and several thousand tracked participants later, the picture has inverted. These memories are not more accurate than ordinary ones. They erode, they drift, and they pick up details that were never there. What makes them unusual is that the erosion is invisible from the inside. Confidence stays high. Vividness stays high. The sense of reliving the moment stays high. Everything that would normally warn a person that a memory has weakened simply fails to arrive [3].
One thing needs saying before anything else. Every number in this article describes groups, not individuals. When a study reports that consistency dropped to sixty-three percent across four hundred people, that is a statement about a distribution. It is not a verdict on any particular reader's memory of any particular day. Some of those memories are close to accurate. The research cannot tell which ones, and neither can the person holding them. That is the whole problem.

The Question Nobody Thought to Ask
In January 1899, a psychologist named F. W. Colegrove published a paper in the American Journal of Psychology with the plain title "Individual Memories" [4]. Abraham Lincoln had been shot thirty-three years earlier. Colegrove asked people what they had been doing when they heard.
They told him. In detail. Where they stood, who came running, what the weather was doing, what they said out loud. Three decades on, the answers arrived with the texture of last week.
Colegrove had no way to check any of it. He simply recorded that the memories existed and that they were unusually rich. And then the observation sat more or less unexamined for the better part of a century.
A note on that citation, because it turns out to matter. The paper is by Colegrove, spelled with an e in the middle, and it appeared in the American Journal of Psychology, volume 10, pages 228 to 255. A surprising number of published papers get this wrong. Talarico and Rubin's own 2003 reference list, in the very study that reset the field, prints the name as "Colgrove" and assigns the paper to the American Psychologist, a journal that did not exist until 1946 [1]. The error has been copied forward for two decades. A field devoted to the study of confident misremembering has been confidently misremembering its own origin.
The modern era began in 1977, when Roger Brown and James Kulik gave the phenomenon a name. They surveyed eighty American men, forty Black and forty white, about a series of assassinations and sudden deaths, plus one shocking personal event of the respondent's own choosing [2]. The John F. Kennedy assassination produced near universal recall. What caught their attention was a split. Memories of Martin Luther King Jr., Malcolm X and Medgar Evers were far more common among Black participants than white ones, which suggested that whatever created these memories was tuned to personal and group relevance rather than to newsworthiness alone.
The specific percentages usually quoted for that split, around seventy-five percent against thirty-three percent for the King assassination, come from teaching summaries and encyclopedia entries rather than from the original tables, which sit behind a publisher paywall. The direction of the effect is not in doubt. The exact figures should be treated as secondhand.
Brown and Kulik then made the claim that would define and eventually sink their theory. Borrowing a neurophysiological idea from Robert Livingston, they proposed a dedicated mechanism. When an event crosses a threshold of surprise and then a threshold of consequence, the brain issues something like a command to print, and lays down a near permanent, unusually complete record of the surrounding moment. They called the resulting memories flashbulb memories, and they described six recurring elements in what people reported: the place, the ongoing activity, the informant who delivered the news, the person's own emotional reaction, other people's reactions, and the immediate aftermath.
Read that list again. Every item is about the moment of hearing. Not one is about the event itself.
This distinction gets lost constantly and it is the single most useful thing to hold onto. A flashbulb memory is not a memory of the assassination or the explosion or the collapse. It is a memory of a hallway, a car radio, a phone call, a face across a kitchen table. The event is what makes the moment worth encoding. The moment is what actually gets encoded.
What Brown and Kulik never did was check whether any of it was true. There was no accuracy test, because there was no ground truth to test against. There was no second measurement, so there was no way to see whether the memories held. There was no comparison to ordinary memories, so there was no way to know whether flashbulb memories were unusual at all. Surprise, importance, emotion and rehearsal were tangled together in a way that made causal claims impossible.
They had documented that people feel certain. They had assumed that certainty tracked accuracy.

Seven Minutes That Broke the Theory
On the morning of January 28, 1986, the Space Shuttle Challenger broke apart seventy-three seconds after launch. The next day, Ulric Neisser handed out a questionnaire to a class of first-year students at Emory University.
Neisser had already spent years arguing that flashbulb memories might be less like snapshots and more like landmarks, stories a person builds and polishes because the event is a personal or historical marker rather than because the brain preserved anything unusual. Now he had a way to test it. He had written accounts collected within a day of the event. All he had to do was wait.
He waited two and a half years. In the autumn of 1988, forty-four of the original respondents answered the same questions again [6].
The results are among the most cited in cognitive psychology, and they deserve a careful framing. Neisser and Harsch scored seven content questions against each person's own original answer, giving a maximum of seven points. The reported mean was 2.95. Eleven participants scored zero, meaning nothing they said matched what they had written. Twenty-two scored two or less. Three achieved a perfect seven. When asked to rate how sure they were on a five point scale, the average came back at 4.17.
Those figures are worth flagging. The chapter itself is a paywalled book contribution and the numbers above circulate through secondary summaries, though they are reproduced consistently in peer reviewed work as well. They should be read as well attested rather than as personally verified in the original.
The direction is not ambiguous. Roughly one in four of these students, all of them reporting high confidence, matched not a single one of the seven details they had written down the morning after the Challenger disaster. Half the sample matched two or fewer.
There was a further detail that turned an interesting result into a disturbing one. In follow-up interviews, participants were shown their own original questionnaires. Their own handwriting. Most did not recover the earlier memory. Some looked at the page, accepted that they must have written it, and continued to report the later version as what they actually remembered. The original had not become inaccessible. It had been replaced.
A pattern also emerged in how the errors moved. Roughly a fifth of respondents said within a day that they had heard about the disaster on television. Two and a half years later, that figure had roughly doubled. Memories were not drifting randomly. They were drifting toward the most available narrative, which was the one people had been watching on screens ever since.
The Challenger study did not stand alone. Michael McCloskey and colleagues ran their own longitudinal study of the same disaster and concluded there was no need to invoke a special mechanism at all [7]. John Bohannon tracked the same event and found a more complicated picture in which emotional intensity and rehearsal pulled in different directions [8]. Sven-Åke Christianson, studying Swedish memories of the assassination of Prime Minister Olof Palme, published a paper whose title said everything: special, but not so special [9]. David Pillemer had already tracked memories of the 1981 assassination attempt on Ronald Reagan and found the same mixture of high confidence and shifting detail [19], and Daniel Wright's study of the Hillsborough football disaster documented errors that were not random but systematically biased in the direction of the dominant public account [18].
What does this mean in practice? It means that the confidence a person feels about a landmark day carries no information about whether the details are right. Not less information than usual. None. Neisser and Harsch found confidence and accuracy unrelated, while confidence and mental imagery were related to each other. Certainty was tracking how clear the picture felt, and the picture kept getting clearer as the memory got further from the truth.
One awkward possibility remained open, though, and critics raised it immediately.

The Control Group That Changed Everything
Here was the objection. Of course the Challenger memories decayed. All memories decay. Nobody had shown that flashbulb memories decayed any faster, or any slower, than the ordinary memories a person forms on an ordinary Tuesday, because nobody had measured ordinary memories alongside them.
Without a comparison, the whole debate was unanchored.
Jennifer Talarico and David Rubin fixed this on September 12, 2001. Their design was simple and it is the reason the study became the field's reference point. Each of the fifty-four participants recorded two memories, not one. The flashbulb memory of hearing about the attacks, and a control memory of an everyday event from within three days beforehand. Same person, same encoding period, same retention interval, same cueing procedure [1].
The retention interval itself was handled with unusual care. Rather than testing the same people repeatedly, which would have turned each test into a rehearsal and contaminated the later measurements, participants were split into three independent groups of eighteen. One group returned after 7 days. One after 42 days. One after 224 days. The intervals were spaced to sit at roughly equal steps on a logarithmic scale, following earlier work on the shape of forgetting curves [13]. Two independent coders scored every detail, agreeing on 96 percent of flashbulb details and 97 percent of everyday details across a subset of reports.
Then came the result that reorganised the field.
Consistency fell. For both memory types. At the same rate. Flashbulb memories lost consistent details and accumulated inconsistent ones on essentially the same trajectory as memories of a forgettable party. There was no interaction between memory type and the passage of time in the recall data at all.
So far, so deflating. But the ratings told a completely different story.
The rating instrument was not improvised. It came from the Autobiographical Memory Questionnaire, built to measure the properties of personal memories rather than their content [12], and the vividness items follow directly from Rubin and Kozin's earlier work on what makes some memories feel unusually clear [5]. Participants rated recollection, meaning the sense of reliving, on a seven point scale. They rated vividness. They rated belief in accuracy, built from whether they believed the event happened the way they remembered it and whether they could be persuaded that their memory was wrong. For the everyday memories, all of these declined over time, exactly as a well calibrated mind should behave. As the memory faded, the sense of having it faded with it.
For the flashbulb memories, none of them declined. Recollection stayed high. Vividness stayed high. Belief stayed high. Ten months on, with the details quietly rearranging themselves, participants felt they were remembering as clearly as they had the day after.
This is the finding, and it needs stating precisely because it is misreported constantly. Flashbulb memories do not decay faster than ordinary memories. They decay at the same rate. What is unusual about them is that the internal warning system goes quiet. The signals a person normally uses to gauge how much to trust a memory keep reading full while the memory itself drains.

There was one more result, and it is the single cleanest number in the whole literature. Talarico and Rubin measured visceral emotional reaction on the day after, asking participants whether their heart raced, whether they felt tense, whether their stomach knotted. They then looked at what that initial physical reaction predicted ten months later.
It predicted belief in accuracy, with a correlation of about 0.31. It predicted nothing about consistency. Not a weak relationship with consistency. No relationship. None of the initial measures, emotional or otherwise, correlated with how accurate the later account turned out to be.
Emotion was building certainty. It was not building accuracy.
Four years later the same authors returned with a longer follow-up and a title that conceded exactly as much as the data allowed: flashbulb memories are special after all, in phenomenology, not accuracy [15]. They are a genuinely distinctive kind of remembering experience. They are not a distinctive kind of record.
A few other findings from that dataset are worth carrying forward. Flashbulb memories were rehearsed far more than everyday ones, 5.23 against 2.46 on the seven point scale. They arrived as more coherent stories and less as fragments, which runs against the intuition that a shocking memory should be splintered. And the question asking for "any other distinctive details" turned out to be responsible for 42 percent of all inconsistencies, in both memory types. The details people volunteer as proof that the memory is real are the ones most likely to be invented.
One subtler result concerned point of view. Ordinary memories tend to drift over time from a first-person perspective toward an outside observer's viewpoint, a shift documented by Georgia Nigro and Ulric Neisser [14]. The everyday control memories did exactly that. The flashbulb memories did not. Participants kept seeing them through their own eyes across the full retention interval, which is one more way the experience of remembering stayed frozen while the content moved.
Charles Weaver had reached a compatible conclusion a decade earlier by asking participants to deliberately encode an ordinary event alongside a shocking one, and finding that confidence in the flashbulb memory outran its consistency in much the same way [11]. His sample was small and the design less controlled, which is why the 2003 study became the citation of record, but the answer was already visible.
Fifty-four students at one university is a small foundation for a claim this large, though. The obvious next question was whether it would hold at scale.
Ten Years, Three Thousand People, One Stubborn Number
It held.
Starting in the week after September 11, 2001, a consortium led by William Hirst and Elizabeth Phelps surveyed more than three thousand people across seven American cities: Boston and Cambridge, New Haven, New York, Washington, St. Louis, Palo Alto and Santa Cruz. Participants were surveyed again at roughly eleven months and roughly thirty-five months. The core longitudinal analyses rest on the 391 people who completed all three waves [16].
The consortium did something the earlier studies had mostly blurred. It separated flashbulb memory, meaning the circumstances of learning, from event memory, meaning the public facts of what happened. These turned out to behave differently, and the difference is instructive.
Flashbulb consistency at eleven months came in at about sixty-three percent. Then the decline nearly stopped. Over the following two years it dropped by only about nine percent proportionally, roughly four and a half percent a year. Statistically the difference between the first and second retention intervals was reliable but small.
Then the consortium came back after ten years [3].
The ten year follow-up added waves at eleven, twenty-five and one hundred nineteen months. Both flashbulb and event memories forgot rapidly during the first year, and then the curves flattened and did not change significantly again, even a decade later. Confidence remained high the entire time.
The bars need a plain explanation, because charts like this are easy to over-read. The first bar is set at one hundred by definition: it is each person's own initial account, the yardstick everything else is measured against. The sixty-three percent at eleven months is reported directly by the consortium. The value at thirty-five months applies the reported proportional decline of about nine percent to that figure. The ten year bar reflects the finding that the curve had flattened and did not shift significantly further, so it is drawn as approximately stable rather than as a separately published statistic.
The flat line is the part that carries the argument. It sits at the confidence level Neisser and Harsch recorded at thirty-two months, 4.17 on a five point scale, which is 83 percent of the maximum. It is drawn as a constant, not as a measured trajectory, because that is what every longitudinal study reports: confidence does not fall the way consistency falls. Hirst and colleagues found it high at every wave including ten years. Talarico and Rubin found belief and vividness unchanged. The line is flat because the finding is that it is flat.
The most unsettling result in the ten year data is not the forgetting. It is what happened to the errors.
Event memories were correctable. When people got a public fact wrong, subsequent media exposure and conversation tended to fix it, and the consortium found that attention to coverage predicted better event memory accuracy. Flashbulb memories behaved in the opposite way. Once an inconsistency entered a person's account of their own morning, it tended to be repeated rather than corrected. A response that was consistent at one wave stayed consistent at the next about eighty-two percent of the time. Errors froze.
There is a reason for this asymmetry, and it is almost obvious once stated. The world will correct a person about how many planes were involved. Nobody will ever correct a person about which friend called them.
Several factors that everyone expected to matter turned out not to. Consistency did not vary reliably with emotional intensity, with how much media someone consumed, with how often they discussed it, with whether they lived in New York, or with whether they had lost someone. After ten years, none of the five measured factors predicted flashbulb consistency at all.
Which raises the question the next section has to answer. If emotion does not protect accuracy, what exactly is it doing?

Before going into mechanism, it helps to see the arc laid out. The field has moved through three distinct phases: a long descriptive period where nobody checked anything, a corrective period built on longitudinal designs, and a current period trying to work out what the phenomenon is actually for.
Notice how long the first gap is. Seventy-eight years passed between the first observation and the first theory, and another fifteen before anyone ran a proper longitudinal test. The phenomenon was described for almost a century before it was measured.
Where the Confidence Actually Comes From
Emotional arousal does improve memory. That part of the original intuition was correct. What it improves, and at what cost, is where the theory went wrong.
The clearest demonstration is chemical. In 1994, Larry Cahill and James McGaugh gave volunteers either a single dose of propranolol, a beta blocker that interferes with the action of noradrenaline, or a placebo, before showing them a slide story. The middle section of the story was emotionally disturbing. The opening and closing sections were mundane. A week later, recall was tested without warning [24].
The placebo group showed the usual pattern: they remembered the emotional middle better than the neutral parts. The propranolol group did not. Their memory for the neutral material was untouched. The emotional advantage had simply been switched off.
Noradrenaline is the brain's version of adrenaline, released under stress and arousal. Blocking its receptors removed the memory boost that emotion normally provides, which means the boost is not a metaphor. It is a specific, interruptible piece of neurochemistry.
The structure doing most of the work is the amygdala, an almond-shaped cluster of neurons buried in each temporal lobe that evaluates emotional significance. James McGaugh's research programme established that it does not store emotional memories itself. It modulates storage elsewhere, principally in the hippocampus, the curled seahorse-shaped structure that binds the elements of an experience into a retrievable episode [23]. Think of the amygdala less as a filing cabinet and more as a supervisor standing over the filing clerk, saying this one, press harder. Reviews of the imaging evidence describe the same division of labour, with amygdala activity at encoding predicting which emotional items survive and hippocampal activity predicting whether they can be bound into a retrievable episode at all [26], and with the two structures interacting rather than operating in sequence [27].
Stress hormones add a further layer, and they are not simply helpful. Glucocorticoids such as cortisol follow an inverted U. Moderate elevation helps consolidation. Very high levels impair retrieval outright, as Dominique de Quervain and colleagues showed by demonstrating that elevated glucocorticoids damaged recall of well learned spatial information [25]. Extreme arousal is not a stronger version of moderate arousal. It is a different regime.

In 2007, Tali Sharot, Elizabeth Phelps and colleagues brought this into contact with September 11 directly. Roughly three years after the attacks, twenty-four people who had been in Manhattan that day were scanned while recalling either their September 11 memory or a control memory from the preceding summer. The group was split in half by where they had physically been. Twelve had been downtown, close to the World Trade Center. Twelve had been in midtown, several kilometres north [28].
The split was sharp. Among the downtown participants, 83 percent showed higher left amygdala activation when recalling September 11 compared with the summer control. Among the midtown participants, only 40 percent did. The interaction was statistically reliable, and the downtown group also described their memories as more vivid and more detailed. Roughly half the total sample reported anything resembling a flashbulb memory, and those who did had been closer to the towers.
What that study established is important, and it is also narrower than it is usually reported to be. Proximity predicted amygdala engagement and subjective vividness. It did not establish that those memories were more accurate. The neural signature tracks the experience of remembering, which is precisely the thing the behavioural literature says is inflated.
There is one more piece, and it explains the specific shape of flashbulb error. Arousal narrows attention. The idea goes back to a 1959 paper by Easterbrook proposing that emotional arousal restricts the range of cues a person uses [29]. Under stress, the centre of an experience is encoded well and the periphery is encoded badly. An older and often overlooked finding sharpens the point: Daniel Reisberg and colleagues showed that it is the quantity of affect, not its quality or its pleasantness, that predicts how vivid a memory feels [32]. Intensity is the variable, and intensity buys vividness rather than fidelity. Christianson's review of the eyewitness literature found the same trade-off repeatedly [30], and the weapon focus studies made it visible: when a weapon is present, eye movements cluster on it and memory for everything else, including the face of the person holding it, suffers [31].
Now map that onto a flashbulb memory. The centre of the experience is the news. The periphery is the room, the informant, the sequence, the clothing, the time of day. The news is what gets protected. The reception context, which is the entire content of a flashbulb memory, is exactly the material that narrowing attention degrades.
Emotion did not build a perfect record. It built a vivid core, a blurred surround, and a strong conviction that both were equally clear.
The Plane That Nobody Saw
On September 11, 2001, two aircraft struck the World Trade Center. Footage of the second impact was broadcast live and repeatedly. Footage of the first impact was not. The only known film of it, shot by a French documentary crew, did not reach television until the following day.
Seven weeks after the attacks, Kathy Pezdek surveyed 690 people in three locations: 275 in Manhattan, 167 in California and 127 in Hawaii. Among other questions, she asked what they had seen on television on the day itself. Seventy-three percent reported that on September 11 they had watched video of the first plane hitting the first tower [33].
That figure gets attributed to all sorts of people. It belongs to Pezdek's 2003 paper in Applied Cognitive Psychology, and it is worth crediting correctly, because it is one of the few places in this literature where a false memory can be checked against a hard external fact.
The most famous single instance involves a president. On at least three occasions in December 2001 and January 2002, George W. Bush described having seen the first plane hit the tower on a television before entering a Florida classroom, and remarked that he had thought it was a terrible piloting error. No such footage existed at that moment. Daniel Greenberg documented the accounts and argued that this is about as clean a naturalistic example of a false flashbulb memory as the field is ever likely to get [34]. The point is not political. It is that heavy rehearsal, constant media exposure and repeated public retelling produced a confidently held memory of something that did not happen, in someone with every reason to have the day well encoded.
The mechanism behind this has a name. Marcia Johnson, Shahin Hashtroudi and D. Stephen Lindsay formalised source monitoring, the process by which a person decides where a piece of remembered information came from [35]. Memories do not arrive tagged with their origin. The brain infers origin from characteristics like perceptual detail and coherence, and that inference can fail. Johnson and Raye's earlier reality monitoring work established the underlying principle: people judge a memory to be real partly on the basis of how much sensory detail it contains [36].
Which sets a trap. Bell and Loftus showed that adding a handful of trivial, unimportant details to a witness account measurably increases how accurate other people judge it to be [37]. Detail is treated as evidence of authenticity, by observers and by the rememberer alike. Television supplies detail in enormous quantities. Over months and years, that footage becomes indistinguishable from personal experience.
Rehearsal makes this worse in a way that is genuinely paradoxical. Talking about the event and thinking about it strengthens the narrative and raises confidence. It also provides repeated opportunities for external material to be absorbed and for the account to be smoothed into a better story. Studies of rehearsal effects have produced inconsistent results precisely because rehearsal does two opposing things at once [38]. The literature on how false memories form describes the same double edge in laboratory conditions.
The digital era has not made this better. A 2022 study of memories of the pandemic outbreak in mainland China examined how dependence on digital media related to personal involvement and flashbulb memory formation, and found media dependence functioning as an active ingredient rather than a neutral channel [39]. The more of an event a person experiences through a screen, the more screen material there is available to be absorbed into what feels like firsthand memory.
What does this mean for anyone trying to assess their own recollection of a public event? Mainly this: the vividness of a televisual detail is not evidence that you were there for it. If a memory of a national tragedy has a camera-like quality, the odds are reasonable that this is because it came from a camera.

Whose Memory Is It
In November 1990, Margaret Thatcher resigned. Martin Conway and a large international team surveyed 369 people and then went back at roughly eleven months. The split they found is one of the cleanest demographic results in the field [40].
Among the 215 UK participants, 85.6 percent met the criteria for a flashbulb memory. Among the 154 participants outside the UK, the figure was 28.6 percent. The difference was enormous and highly reliable. A subset of thirty-three Lancaster students was followed to twenty-six months, and thirty-one still had their flashbulb memories intact.
One correction is worth making here, since it circulates in summaries. The non-UK sample was not an even mix of Danish and North American respondents. It was roughly ninety-five percent North American with a small Danish minority.
The interpretation is straightforward. The same event, the same news coverage, the same day. What differed was whether it mattered to the person hearing it. Relevance, not drama, does the work.
Age matters too, and the picture is more complicated than it first appeared. Gillian Cohen, Martin Conway and Elizabeth Maylor found that ninety percent of younger adults met the strict flashbulb criterion for the Thatcher resignation, against forty-two percent of older adults [41]. That looks decisive until you notice that the same study found no comparable age gap for the Hillsborough football disaster.
A meta-analysis settled the general question in 2020. Sarah Kopp, Laura Sockol and Kristi Multhaup pooled sixteen studies covering 1,898 participants, comparing adults under forty with adults over sixty [42]. After removing outliers, the effect on flashbulb memory scores was a Hedges' g of -0.30 across fourteen studies, with a confidence interval from -0.45 to -0.15. Consistency showed a similar effect of -0.29 across seven studies.
That is a small to moderate impairment, not a collapse. Older adults form flashbulb memories readily. They form slightly fewer of them, and slightly less consistently. This is one place where the popular summary and the pooled evidence diverge noticeably, and the citation is also frequently mangled: the lead author is Kopp, not Koppel, which is a different researcher working in the same area.
A related finding suggests the age effect may be about something other than aging as such. Patrick Davidson and Elizabeth Glisky asked whether flashbulb memory is really a special case of source memory, and found that among older adults, the ability to form these memories was unrelated to measures of frontal lobe function [43].
The relevance principle from the Thatcher study has been tested more directly since. Talarico, Bohn and Wessel examined memories of the Fukushima nuclear disaster across three countries with different relationships to nuclear power. Across samples of 265 and 518 participants, German respondents were the most likely to hold flashbulb memories of it, matching the prediction that events congruent with a group's existing beliefs and concerns are the ones that stick [44]. Earlier cross-national work on the death of French President François Mitterrand found the same shape [45].
Flashbulb memories, on this reading, are not simply records of shocks. They are markers of belonging. They encode the moment a person discovered that something had happened to a group they consider themselves part of.
Two boundary cases test how deep the phenomenon runs. Dorthe Berntsen and Dorthe Thomsen studied Danish memories of the German occupation and liberation more than half a century later, and found that people who had been active in the resistance had more accurate and clearer memories than those who had not [46]. And a recent study of flashbulb memories across the lifespan in Alzheimer's disease found that while access to these memories was reduced, the ones that survived clustered in the same reminiscence period and retained the same canonical structure as in healthy older adults [47].
Even as the machinery degrades, the shape of the phenomenon holds.

Four Theories and One Survivor
The field produced four serious models. Comparing them side by side shows how the explanation migrated away from biology and toward meaning.
Two of these deserve a note on method, because they were tested in an unusual way. Catrin Finkenauer and colleagues used structural equation modelling on 394 Belgian respondents remembering the death of King Baudouin, fitting competing causal paths against the data and comparing how well each reproduced the observed correlations [48]. Nurhan Er applied the same approach to 655 people after the 1999 Marmara earthquake in Turkey, with the crucial difference that 335 of them had been direct victims [49]. In that group, the distinction between remembering the event and remembering hearing about it collapsed, because they had not heard about it. They had been in it.
The trajectory across these four models is a steady retreat from biology. Brown and Kulik proposed a dedicated brain mechanism. Conway proposed a combination of ordinary factors. Finkenauer proposed a causal chain running through emotion and rehearsal. Er proposed that personal importance is the engine and emotion the transmission. By 2016, Hirst and Phelps summarised the consensus position bluntly: flashbulb memories do not require special memory mechanisms, and are best characterised as involving both forgetting and distortion despite high confidence [22].
The flashbulb is not a different kind of bulb. It is the same bulb, pointed at something that mattered.
A Pandemic Is Not a Gunshot
Everything above was built on sudden events. Assassinations, explosions, attacks. Brown and Kulik made surprise a requirement, and every subsequent model kept it in some form.
Then came an event with no moment.
The COVID-19 pandemic unfolded over weeks and months. There was no single detonation. If surprise is genuinely necessary, a slow-onset global crisis should not produce flashbulb memories at all. It should produce ordinary, gradually accumulated autobiographical knowledge.
It did not.
The largest test came from a cross-national survey published in Memory in 2024, covering eleven countries and examining memory for learning about the first COVID-19 case in each nation [50]. Participants held detailed memories of the date and of who else was present, while memories of place, ongoing activity and news source were only partially detailed. Specificity was highest in China. A classification and regression tree analysis identified the strongest predictors: age, how severe the person judged the situation to be, whether they lived somewhere with stringent protective measures, and what they expected the pandemic to do.
A second study tracked American students across the closure of their campus, with surveys in March and May 2020 [51]. Flashbulb memories formed. They were more consistent than the participants' own control memories, and confidence did not differ between the two types. The lead author on this one is Xuan, which is worth stating because the name is sometimes rendered incorrectly in secondary discussions.
Spanish researchers studied memory for the declaration of the national state of alarm across three age groups and found the expected markers of a flashbulb memory: detailed recall, strong reliving, high confidence. Younger adults recalled canonical categories in more detail than the other groups, and middle-aged adults patterned with the older group rather than the younger one [52]. Separate work found that the pandemic reorganised how people date their own autobiographical events, functioning as a boundary in personal time [53].
Something else happened to pandemic memory that has no parallel in the September 11 literature. A study published in Nature in 2023 found that retrospective accounts of the pandemic were systematically distorted in self-serving directions, with people's recollections of their own past risk assessments and attitudes bending to fit their current position [54]. Memory of a prolonged, politically contested event does not just fade. It gets edited toward the version that is currently comfortable.
So what happened to the surprise requirement?
It has been demoted rather than discarded. The resolution the field has converged on is to separate the protracted event from the punctate moment of reception. A pandemic has no single moment, but learning that the first case had arrived in your country does, and so does an announcement that your campus is closing tomorrow. The reception moment can be sharp even when the event is not.
Evidence from other predictable events points the same way. Analysis of the 2016 UK referendum, tracking 851 participants across four time points over sixteen months, examined an outcome that was anticipated, campaigned over for months, and emotionally opposite for different voters. Memory differences emerged over time along the lines of emotional valence rather than surprise [56]. A study of the 2016 American election night found higher memory confidence among one group of voters and higher rehearsal among the other, despite similar levels of actual information quantity and consistency [57]. And work on unexpected positive events found that pleasant surprises do not reliably generate flashbulb memories at all [58].
A study of the January 2021 Capitol riots, comparing seventy-nine American with seventy Belgian participants, added a further wrinkle. Americans reported more flashbulb memories, as relevance would predict, but the Belgians' accounts of the event itself and its causes were more similar to one another than the Americans' were [55]. Distance from an event can produce a cleaner shared story and a weaker personal one at the same time.
Surprise, it turns out, is neither sufficient nor strictly necessary. Consequence is doing the heavy lifting.

What the Field Still Argues About
An honest account has to say where the ground is soft.
The rate of long-term forgetting is genuinely disputed. The consortium data show a steep first year followed by a plateau that holds for a decade [3]. Schmolck, Buffalo and Squire's study of memories of the Simpson verdict found the opposite shape, with distortion accelerating between fifteen and thirty-two months and the proportion of highly accurate accounts falling from about half to under a third [17]. The consortium has argued that interference from a subsequent civil trial explains part of the difference. That explanation is plausible and it is not settled.
Consistency estimates for the same event vary more than they should. Another large-scale analysis of September 11 memories reported figures around eighty percent at eleven months [20], against the consortium's sixty-three percent, and further work on very long delays has produced its own picture [21]. Same event, same country, overlapping years, different numbers. Coding rules and question wording account for much of the spread, which is itself a finding about how fragile these measurements are.
Rehearsal remains unresolved. Some studies find it improves consistency, some find no relationship, and some find it associated with more error. Given that rehearsal simultaneously strengthens a narrative and imports foreign material into it, the mixed results may be the correct answer rather than a failure to find one.
The methodology deserves a harder look than it usually gets. Nearly every study in this article measures consistency, not accuracy. Consistency means agreement between a person's later account and their own earlier account. Talarico and Rubin stated the limitation precisely: consistency is measurable and is a necessary though not sufficient condition for accuracy, and an inconsistency implies that at least one of the two reports is wrong [1]. But a memory can be perfectly consistent and consistently false. The Challenger cueing interviews and the Pezdek first-plane data are among the few places where an external fact was available to check against.
Timing of the first measurement is a hidden confound. Winningham, Hyman and Dinnel showed that when the initial report is collected matters a great deal, because a baseline taken a week after the event already contains a week of rehearsal and contamination [10]. Studies that started late are effectively measuring the stability of an already altered memory.
And the act of measuring changes the thing measured. Asking someone to write out their memory is itself a rehearsal, one that can consolidate and reshape the account. This is why Talarico and Rubin's use of three independent groups rather than repeated testing was a real methodological advance, and why studies that test the same people again and again should be read with that caveat attached. Anyone who has read about how memories are rewritten during reconsolidation will recognise the problem.
Finally, a minority position deserves acknowledgement. A recent review in Nature Reviews Psychology takes the malleability of emotional autobiographical memory as its starting point while treating the persistence side of the phenomenon as equally real and equally in need of explanation [62]. Nobody now defends a literal photographic mechanism. Some researchers do continue to argue that emotion plays a genuinely distinctive role in formation, even if not in preserving accuracy.
The Boundary Around Confidence
There is a temptation to walk away from this literature with a simple slogan: confidence is worthless. That would be wrong, and getting it wrong matters, because the same claim shows up in courtrooms.
In 2017, John Wixted and Gary Wells published a synthesis of the eyewitness confidence and accuracy literature that reached an apparently opposite conclusion [61]. Under what they call pristine testing conditions, confidence is a strong indicator of accuracy. Pristine means a fair lineup with only one suspect, unbiased instructions, a double-blind administrator, and a confidence statement taken at the moment of the initial identification, before any feedback or discussion.
Under those conditions, a highly confident witness is usually right.
This does not contradict the flashbulb findings. It locates them. The flashbulb literature measures confidence months or years after the fact, in memories that have been rehearsed dozens of times, discussed with everyone the person knows, and saturated with televised imagery. That is the exact opposite of pristine.
The two literatures agree on the underlying principle, which is more useful than either alone. Confidence carries real information at the moment of an uncontaminated first retrieval, and it stops carrying information as delay and contamination accumulate. What makes flashbulb memories dangerous is that the confidence does not decay along with the diagnostic value. It stays at full strength long after it has stopped meaning anything.
A useful way to hold both facts: confidence is a reading taken at a specific moment, and like any reading it has an expiry date. The flashbulb finding is that nobody feels the expiry date pass.

The Lesson That Transfers
Almost nobody reading this will be asked to testify about a national tragedy. The reason the flashbulb literature is worth knowing has less to do with public events than with a mechanism that operates every day, in much smaller ways.
Here is the mechanism, stripped to its core. The mind judges how well it knows something by consulting how the knowing feels. Clear, fluent, vivid, easy: must be solid. Effortful, patchy, slow: must be shaky. That heuristic works acceptably most of the time, which is why it survives. Flashbulb memories are the case where it fails completely and visibly.
Metamemory research arrived at the same conclusion from a different direction. Asher Koriat showed that judgments of learning, the predictions people make about what they will later remember, are built from cues available at the moment of judging, and that those cues can point the wrong way [59]. Aaron Benjamin, Robert Bjork and Bennett Schwartz went further, demonstrating that retrieval fluency, meaning how quickly and smoothly something comes to mind, predicts confidence while being dissociated from actual later recall [60]. Their title said it plainly: the mismeasure of memory.
Now consider what happens during exam preparation. A student rereads a chapter. The third pass feels smooth. Sentences arrive before the eye reaches them. Nothing surprises. That fluency gets read as mastery, and the student stops.
Then the exam asks for the material without the page in front of them, and it is not there.
The flashbulb literature is the extreme version of this same error, run on real people over real decades with real stakes. It demonstrates that vividness and confidence can stay locked at maximum while the underlying content quietly degrades, and that the person holding the memory receives no signal at all. If that can happen with the most emotionally charged day of someone's life, it can certainly happen with a chapter on cellular respiration. Work on the illusion of knowing describes the same failure in study contexts.
The practical implication follows directly and it is not complicated. Subjective confidence is not a measurement. It is a feeling that correlates with measurement under good conditions and stops correlating under bad ones. The only way to find out whether something is actually retrievable is to try to retrieve it, without the source available, and see what comes back. Everything else is a guess dressed up as a certainty.
There is a second implication that gets less attention. Talarico and Rubin found that visceral emotional reaction predicted later belief in accuracy with no corresponding effect on consistency, and separately predicted later symptoms of post-traumatic stress. Emotion is doing something real to memory. It is just not doing the thing people assume. Understanding how emotions shape memory means accepting that arousal reliably strengthens the conviction and only unreliably strengthens the content.
Belief and recollection are separable constructs, and researchers have spent considerable effort pulling them apart [63]. A person can believe an event occurred without vividly reliving it, and can vividly relive something that did not occur as remembered. Flashbulb memories sit at the far corner of that space: maximum reliving, maximum belief, unremarkable accuracy.
Conclusion
The strangest thing about this body of research is what it did not find.
It did not find that flashbulb memories are fragile. They are about as durable as anything else in autobiographical memory, and after the first year they are remarkably stable. It did not find that they decay faster than ordinary memories. Talarico and Rubin's control condition settled that question in the opposite direction. It did not even find that they are especially inaccurate in absolute terms. Sixty-three percent consistency at eleven months is not a catastrophe.
What it found is that the relationship between how a memory feels and how good it is comes apart, and stays apart, for decades.
That is a stranger result than simple unreliability, and a more useful one. A memory that failed loudly would be manageable. People would notice the gaps and hedge accordingly. What the evidence describes instead is a memory that fails silently, with every internal indicator reading normal. The certainty is not a symptom of accuracy. It is a separate output of the same emotional machinery, produced in parallel, and it does not know what the other output is doing.
Colegrove noticed the vividness in 1899 and could not test it. Brown and Kulik named it in 1977 and mistook the vividness for evidence. It took the Challenger disaster, an everyday memory control condition, three thousand people across seven cities, and ten years of follow-up to establish what should perhaps have been suspected from the start.
The flash was never the record. The flash was only ever the feeling of having one.
Frequently Asked Questions
Are flashbulb memories more accurate than normal memories?
No. When Talarico and Rubin compared September 11 memories against everyday memories from the same period in the same people, consistency declined at the same rate for both. What differed was that vividness, reliving and belief in accuracy stayed high only for the flashbulb memories.
Does a vivid memory mean it definitely happened that way?
Vividness reflects how much perceptual detail a memory contains, not how well it matches events. Because the mind treats detail as evidence of authenticity, imported material from television or repeated retelling can make an inaccurate memory feel more real rather than less.
Why do people remember the day but not the details correctly?
Emotional arousal narrows attention onto the central content and away from the surround. The news itself is encoded well. The room, the informant and the sequence, which are the actual substance of a flashbulb memory, are exactly what narrowed attention degrades.
Do older adults form flashbulb memories?
Yes, readily. A meta-analysis of sixteen studies covering 1,898 people found a small to moderate age difference, with a pooled effect size of about -0.30 for flashbulb memory scores. That is a modest reduction in how often and how consistently they form, not an absence.
Did COVID-19 produce flashbulb memories despite being a slow event?
It did, which forced a revision. Researchers now separate the protracted event from the sharp moment of reception, such as learning of a first national case or a campus closure. Surprise appears to help these memories form rather than being strictly required.




