The Grocery List Is the Worst Example Anyone Ever Chose

You know this one already, even if the name is new. Somebody reads you ten things to pick up. You get milk, because it was first. You get batteries, because it was last. Somewhere between them was a fourth item, and it is gone.

Psychologists call the shape of that failure the serial position effect. Items at the start of a list get recalled well, items at the end get recalled well, and the ones in the middle fall into a trough.

Plot recall against position and you get a U with a long flat bottom.

The grocery list is where every account of this starts, and usually where it stops. Which is a shame. It is the least interesting thing this curve has been measured on.

It also turns up in rats in a maze, in people who cannot form new memories, in expert judges scoring a cooking competition, and in language models reading twenty documents at once.

Here is the part that gets left out. The explanation you were probably taught, where the start of the list is long-term memory and the end is short-term memory, has been in open dispute since 1974.

Not gently qualified. Disputed, with experiments on both sides, for fifty years.

And the middle, the part treated as the leftover in every retelling, turns out to be the part that can be dissociated from the rest. There is a drug that leaves both ends of the curve untouched and impairs only the middle.

The middle does not fade first, either. It starts low and stays low.

So: the curve and the numbers behind it, the account you were given and what broke it, and the question that usually goes unasked. Does any of this happen outside a laboratory word list?

1957 and 1962: The Shape Before Anyone Could Explain It

The curve arrived before the theory did, which is usually a good sign.

James Deese and Roger Kaufman reported in 1957 that when people freely recall a list, the serial position of an item predicts whether it survives, and that the pattern changes depending on whether the material has its own internal order [1]. The shape was there. Nobody yet had a theory that explained it. The definitive description came five years later. Bennet Murdock ran 103 students through six conditions, varying list length across 10, 15, 20, 30 and 40 words and presentation rate across one or two seconds per word [2]. What he described is still the standard picture: a steep primacy effect covering roughly the first three or four words, an S-shaped recency effect covering roughly the last eight, and between them a horizontal asymptote.

That word asymptote is worth pausing on. It does not mean the middle is empty. It means the middle is flat.

Between the fourth item and the eighth from last, an item's chance of being recalled is about the same as its neighbours. The curve drops, runs level for a long stretch, and rises.

A flat region is a strange thing for a memory to produce. If forgetting were only a matter of time passing, position eight and position fifteen should differ, because one waited longer. They do not differ much. Something keeps the middle at a constant low value, and what that something is turns out to be three things.

Edward Feigenbaum and Herbert Simon published a theory of the effect in 1962, the same year as Murdock's data, arguing that the curve falls out of the order in which a learner allocates limited processing to items rather than out of any separate storage system [3]. That idea went quiet for decades. Worth knowing it was there in 1962, because a version of it is winning now.

The curve also shows up when you ask people to recall the list in order rather than in any order they like. John Jahnke reported serial position effects in immediate serial recall in 1963, with primacy dominating and recency compressed into the final position or two [4]. Same shape, different task, different balance between the two ends.

1957
Deese and Kaufman report that list position predicts recall
1962
Murdock maps the full curve across 103 people
1966
Glanzer and Cunitz erase recency with 30 seconds of counting
1971
Rundus counts rehearsals aloud: four or five per early item
1974
Bjork and Whitten distract after every item and recency returns
1990
Koppenaal and Glanzer call long-term recency an artefact
1997
Nairne states the ratio rule for recency
2002
Howard and Kahana publish a distributed representation of temporal context
2007
Brown Neath and Chater formalise distinctiveness as SIMPLE
2012
Serruya records the list from inside 84 human brains
2024
Liu finds the same U shape in language models
2025
Blokland finds a drug that impairs only the middle
2026
Polyn and Woodman drop the capacity limit entirely

The Experiment That Settled It, and What It Actually Showed

In 1966 Murray Glanzer and Anita Cunitz published the study that turned a curve into a doctrine.

Their first experiment gave 240 men lists of 20 words each, varying presentation rate and the number of times each word appeared [5]. Repeating a word helped the early and middle positions. It did nothing for the last few. That is already a dissociation.

The second experiment is the famous one, and it is smaller than people assume. Forty-six men learned 15-word lists [5]. After the last word they either recalled immediately, or counted aloud from a randomly chosen digit for 10 seconds, or counted for 30 seconds [5]. After 10 seconds most of the recency peak was gone. After 30 seconds there was nothing left of it. The primacy end did not move.

Leo Postman and Laura Phillips had reported the same delay result in 1965, a year earlier and working independently [6]. Two laboratories, one conclusion. Nothing since has overturned it.

The interpretation is what came under fire. Two storage mechanisms, Glanzer and Cunitz argued.

Recency is the contents of a short-term store being read out, and counting backwards fills that store with digits and pushes the words out. Primacy is what made it into long-term memory, and no amount of counting touches that.

Read as a description of the data, that is fair. Read as a claim about the architecture of memory, it is a much bigger leap, and it is the leap that ended up in every textbook.

You can see why it was irresistible. It explained a strange shape with two familiar things, and it made a prediction that came out right the first time. Good theories do that. So do wrong ones.

Counting the Rehearsals Out Loud

If primacy is about getting words into durable storage, the obvious question is what the person is doing that gets them there. In 1971 Dewey Rundus answered it by asking people to rehearse out loud and counting what they said [7]. The method had been introduced in 1970 with Richard Atkinson [8]. Say whatever you are repeating to yourself, into a microphone, while the list plays. Crude, and it produced one of the tidiest results in the field.

In that 1971 count, early items got about four or five rehearsals each and late items got one or two [7]. And the recall probability for the early part of the list tracked the rehearsal count closely. The first word has nothing to compete with, so it gets repeated alone. The second word gets repeated with the first. Well before the end of the list the rehearsal set is full, and each new item is lucky to be said at all.

Rehearsal alone does not make a memory durable, though, and Fergus Craik and Michael Watkins showed that in 1973 by holding items in mind for varying lengths of time without any deeper processing [9]. Time spent repeating a word did not predict later recall of it. What mattered was what you did with it. That looks like a contradiction with Rundus. It is not quite one. A rehearsal count predicts recall because each pass is another chance to encode, not because repetition on its own deposits anything. Which is one reason the field still argues about what primacy is. That argument is the difference between rote repetition and elaborative rehearsal.

There is a further wrinkle. In 1977 Delbert Brodie and Bennet Murdock plotted the curve by the order in which items were actually rehearsed rather than the order they were presented in, and the primacy advantage shrank considerably [10]. Position one is not magic. It is where the rehearsals pile up.

1974: One Distractor in the Wrong Place

Here is the result that changed everything.

Robert Bjork and William Whitten reasoned that if recency is a short-term store being emptied, then filling that store during the list should destroy recency just as thoroughly as filling it after the list. So in 1974 they put a distractor task after every single item, including the last one [11].

By the Glanzer and Cunitz logic there should have been no recency effect at all. The store had been emptied item by item all the way through, and once more before recall.

The recency effect came back.

That finding is called long-term recency, and it is the hinge the whole field turns on. Steven Poltrock and Colin MacLeod confirmed both primacy and recency in the continuous distractor paradigm in 1977 [12]. The advantage for the last few items survives conditions under which no short-term store could still be holding them.

Take a second with that. The delay that kills recency in the standard task is 30 seconds of counting. In the continuous distractor task the last item gets the same counting and is still recalled better than the middle. The difference is not whether there was a delay. It is what happened before it.

That is the shape of a timing effect, not a storage effect. And once you see it as timing, a different explanation becomes available.

The Ratio That Predicts Recency Better Than a Store Does

Look down a row of telephone poles. The nearest two are easy to tell apart. The far ones blur together, not because they are smaller, but because the ratio of the gap between them to their distance from you has collapsed.

Temporal distinctiveness accounts say memory works that way, looking backwards through time. The last item is recent and the one before it is only slightly less recent, but the ratio between those two gaps is large when you look immediately and small when you look after a delay.

Robert Greene reviewed the recency literature in Psychological Bulletin in 1986 and argued that immediate and delayed recency effects share a single basis rather than reflecting two different systems [13]. Arthur Glenberg and Naomi Swanson set out the temporal distinctiveness version in 1986 too, showing it also handles the fact that spoken lists produce bigger recency than written ones [14].

James Nairne and colleagues stated the rule in its sharpest form in 1997. What predicts the size of the recency effect is the ratio of the interval between items to the interval before recall [15]. Widen the gap between the items to 30 seconds and recency survives a delay that would otherwise erase it [15].

Not the contents of a store. A ratio between two durations.

One idea, both results. In the standard task the items arrive a second apart and then you count for 30 seconds, so the ratio is terrible and recency dies.

In the continuous distractor task the items are themselves 30 seconds apart, the ratio is roughly preserved, and recency lives. No store required.

Gordon Brown, Ian Neath and Nick Chater formalised this as a temporal ratio model in Psychological Review in 2007, and it reproduces serial position curves across a range of tasks with no short-term store in it [16]. Ian Neath had already shown in 1993 that distinctiveness predicts serial position effects even in recognition, where the classic explanation has less to say [17].

A second family of accounts reaches a similar place by another route. Marc Howard and Michael Kahana argued in 1999 that what you retrieve is not an item but a slowly drifting context, and that the curve falls out of how much that context has changed between study and test [18]. They formalised it as a distributed representation of temporal context in 2002 [19], and Per Sederberg and colleagues extended it into a context-based theory covering both recency and the tendency to recall neighbouring items together [20].

Two replacement theories, and they disagree with each other as well as with the original. That is what a live field looks like.

The Two-Store Side Did Not Concede

It would be dishonest to leave it there.

Lois Koppenaal and Murray Glanzer came back in 1990 with three experiments arguing that long-term recency is not what it looks like. Their claim was that participants adapt to a distractor that repeats throughout the list, learn to time-share with it, and thereby keep short-term storage working, an argument they made across three experiments [21]. In their words, the long-term recency effect is "really a short-term storage effect, resulting from adaptation to the repeated presentation of a particular type of distractor throughout the list". Their first experiment showed that an appropriate post-list distractor, one the participant has not adapted to, eliminates the long-term recency effect entirely.

That is a serious answer, and if it holds, the 1974 result does not show what it appears to show.

Imaging has been brought in on both sides of this dispute. Deborah Talmi and colleagues scanned 10 young adults recognising items from early or late positions and found that recognition of early items uniquely engaged regions associated with long-term retrieval, which is the dual-store prediction [22]. Ten people is a small study and the authors present it as evidence rather than proof. The same researcher later produced a result pointing the other way. In 2015 the group tested nine patients with anterograde amnesia against 15 control participants [23]. By the classic account, patients with impaired long-term memory should show recency and nothing else. The continual distractor task produced long-term recency in them anyway, and their advantage for recent items over middle items was as large as the controls' [23]. That study is small too, as patient studies are. Single-store accounts predict the pattern. Dual-store accounts have to work to accommodate it.

The exchanges got specific after that. In 2008 Simon Farrell and Stephan Lewandowsky published limits on how far a retrieved-context account can stretch to fit the data on recalling neighbouring items [24]. Howard replied in the same journal in 2009, showing the interactions they described were predicted by the temporal context model after all [25]. Bennet Murdock, whose 1962 data started the whole thing, published a formal objection to the SIMPLE model in Psychological Review in 2008 [26].

So where does that leave you? Not with a winner. The honest summary is that the U-shaped curve is one of the most reliable findings in psychology, and that the field has spent five decades arguing about which half of it needs its own memory system, if either does.

The newest shot comes from an unexpected direction. Sean Polyn and Geoffrey Woodman showed in 2026 that several of the signature results used to prove you are measuring a limited-capacity working memory store fall out of a long-term memory model that has no capacity limit in it at all [27]. The diagnostic tests, in other words, may not be diagnostic.

What Is Actually Wrong With the Middle

Now the middle itself.

Every popular account defines the middle by subtraction. The start gets rehearsal, the end gets recency, the middle gets neither. That is not a mechanism.

It is an accounting identity, and it is wrong, because three separate things go wrong for a middle item.

Start with rehearsal, which arrives first and sets up the other two. Once the rehearsal set is saturated, the 1971 counts show each new item getting one or two passes instead of the four or five the opening word received [7]. The middle is not unrehearsed. It is thinly rehearsed, and thin rehearsal is the maintenance kind that Craik and Watkins showed in 1973 does not build a durable trace on its own [9]. So the middle item enters the competition already weak, which is what makes the next two costs bite.

That competition is what interference describes, and the middle is the only stretch of a list exposed to it from both sides.

The opening item has nothing before it. The closing item has nothing after it.

Everything between is squeezed by earlier items competing at retrieval and later items overwriting the context it was encoded in, which is the double bind that proactive and retroactive interference describe between them.

Timing then removes the last escape route. A weak trace can still be found if something about when it happened marks it out, and under the 1997 ratio account an item in the middle sits among neighbours at almost identical temporal distance, so its position carries no cue at all [15].

The telephone poles blur. Thin encoding, competition on both sides, and no timestamp worth having, all landing on the same items.

Item in the middle

Rehearsal set saturated

Neighbours on both sides

No temporal edge

One pass not five

Interference from both directions

Position carries no cue

Weak trace at recall

Three costs, one region. And the region itself can be taken out on its own, which is the strongest reason to think it is a place where something happens rather than a place where nothing does.

In 2025 Arjan Blokland and colleagues pooled data from four double-blind randomised studies of biperiden, a drug that blocks M1 muscarinic receptors and is used as a pharmacological model of the memory loss seen in early Alzheimer's disease [28]. They expected it to flatten primacy, because that is what the disease does. It did not. Biperiden left both primacy and recency intact and impaired memory for the middle 10 words, the study scoring each list as its first 3 words, its last 3, and the 10 between them.

Read that again.

A drug damaged the middle of the curve and left the ends alone. Whatever the middle is, it is not simply the absence of the two things that help the ends. Something specific is happening there, and a drug can interfere with it while leaving the ends alone.

There is an older result pointing the same way. Isolating one item in a list, by printing it in a different colour or making it stand out in any way, rescues it, and a 1966 study found the size of that rescue depends on where in the list the isolated item sits [29]. You can pull an item out of the trough. Distinctiveness is the lever.

The Sound of the Last Word

Something odd happens when a list is spoken rather than shown, and the explanation offered for it added a third store to the two already in play.

Spoken lists produce a bigger recency effect than written ones. Bennet Murdock and Keith Walker documented the modality effect in free recall in 1969 [30], and Conrad and Hull had shown the input modality changing the shape of the serial position curve the year before [31].

Hearing the list is better for the end of it than reading the list is.

Then comes the strange part. Play one extra sound after the final word, a sound the person knows carries no information, and the auditory advantage disappears. Robert Crowder and John Morton proposed in 1969 a brief precategorical acoustic store to explain it [32], and Morton ran the experiments that mapped the effect out in 1971 [33]. One irrelevant syllable, spoken after you have finished, and the last word of the list is worse off.

You have felt this. Somebody gives you a phone number and then asks whether you got it. The question is not the problem. The sound is.

The neat acoustic explanation did not survive intact. In 1982 Pennie Ottley and colleagues showed the suffix effect depends on how the listener interprets the extra sound rather than on its acoustics alone [34], which pushes it up out of the ear and into perception. And in 1985 Rena Krakow and Vicki Hanson found serial recall effects in deaf signers working in the visual modality [35], which is difficult to explain with a store built specifically for sound.

The modern picture is more balanced than either. Rachel Grenfell-Essam and colleagues tested lists ranging from 2 to 12 words in 2017 and found an auditory advantage at the end of the list and a visual advantage early in it, with the size of both depending on list length [36]. Reading is better for the beginning. Hearing is better for the end.

If you want the underlying machinery, the very short buffers that these effects live in are covered in the article on sensory memory.

What the Brain Does Across a List

You can now watch this happen from the inside, and the recordings say something the behavioural work could not.

Serruya and colleagues analysed intracranial recordings from 84 patients undergoing neurosurgery in 2012 as they studied and recalled word lists [37]. Gamma power between 28 and 100 Hz was high at the start of the list and gradually subsided across positions, while lower frequency delta and theta power between 2 and 8 Hz did the opposite. The shift was widespread rather than local, and the position effect on encoding success was strongest in medial temporal lobe structures.

That is a brain state that changes as a list unfolds. Not a switch between two stores. A drift.

Ryan Colyer and Michael Kahana added a specific piece in 2025. Across 188 patients studying unrelated words, higher theta phase consistency correlated with better recall of early list items in particular [38]. It held again in a second group of 157 patients studying categorised words [38].

Attention is part of the story. Jacqueline Kim and colleagues recorded EEG in 2026 while people encoded lists either with full attention or while doing a target detection task at the same time, and found that dividing attention flattened the normal gamma decline across serial positions [39]. If you want the general version of that argument, it is in the piece on attention and memory.

These signals now predict individual trials. Yuxuan Li and colleagues had 98 participants recall 576 lists across 24 sessions, then trained classifiers on the EEG recorded just before each word was spoken. Those classifiers could tell whether the spoken word belonged to the target list, and classifiers trained on activity during encoding predicted which words would later be recalled [40]. Sarah Seger and colleagues used direct recordings from hippocampus, orbitofrontal cortex and posterior cingulate simultaneously in 2025, and found the relative timing of ripple events across those regions tracked successful encoding of order [41].

The animal work is older and blunter and makes the same point. Raymond Kesner and Jeanne Novak trained rats on eight-item lists in a radial maze in 1982, and normal rats showed both primacy and recency. Rats with dorsal hippocampal lesions lost primacy and kept recency, and a delay of 10 minutes before the test impaired every part of the curve [42]. In 2005 the same laboratory narrowed it further, showing that CA3 lesions specifically disrupted performance on the first three serial positions [43]. Terrace showed in 1993 that pigeons and monkeys learn serial lists too [44]. A rat has no language and no revision technique, and neither has a pigeon.

Whatever produces this curve sits well below anything you would call studying. It is old.

When the Curve Becomes a Clinical Signal

The dissociation that shows up in lesioned rats also shows up in people, and it has become a measurement.

Daniel Weitzner and Matthew Calamia reviewed the literature on serial position effects in mild cognitive impairment and Alzheimer's disease in 2020 [45]. Across studies, both groups showed a similar pattern of reduced primacy with relatively intact recency, and the loss of primacy was more pronounced in Alzheimer's disease. Their review also found that these position measures predicted future decline, in some cases beyond what the overall memory score predicted.

That last clause is why anyone cares. A total score says how much someone remembered. The shape of the curve says something the total does not.

Two 2025 studies put numbers on it. Veronica Di Palma and colleagues followed 62 patients with mild cognitive impairment for three years [46]. Thirty patients progressed to Alzheimer's disease and 32 patients stayed stable, and delayed recall of the primacy portion of a story predicted which group a patient ended up in [46]. Not a word list. A story, which is closer to what a person actually encounters. Working from the other end of the curve, Leonardi and colleagues compared 54 people reporting subjective cognitive decline against 69 healthy controls on a 15-word list [47]. Both groups showed a normal recency effect immediately. The difference appeared after a delay, as a group by position interaction with F = 4.47 and p = .01, driven by faster forgetting from the recency portion in the group with complaints.

The pattern is not universal across conditions. Paul Massman and colleagues compared Parkinson's and Huntington's disease patients in 1990 and found many shared deficits, but abnormal serial position recall appeared in the Huntington's group and not the Parkinson's group [48].

Different diseases bend the curve differently, or not at all.

This is research about groups. Be plain about what it is not.

None of it is a self-test. A curve from one person on one list carries no diagnostic information, and reading your own recall as a signal about your brain misuses every study in this section.

Does Any of This Survive Outside a Word List?

Here is the question, and the answer has two halves that point in opposite directions.

The first half is that yes, the effect appears with real material. Kelly Bennion and colleagues had people encode and recall lists of short TikTok videos in 2025 rather than words, and primacy, recency and proactive interference across lists all survived the switch [49]. Video is not a word list. The curve did not care. It appears in judgement too. Maira Emy Reimão and colleagues examined scoring on a cooking competition show in 2025 and found that where a dish came in the tasting order predicted the score it was given, among expert judges with real stakes on the outcome [50]. Order effects in persuasion go back further, to Norman Miller and Donald Campbell in 1959, who showed that whether the first or the last argument wins depends on the timing of the speeches and of the measurement [51]. So much for going last.

Now the second half, which is the one that should make you cautious.

Benjamin Rottman and Yiwen Zhang had people learn a cause-effect relationship that changed over time, once in the usual fast laboratory version and once spread across 24 days on a mobile phone [52]. Over 24 days, summary judgements showed a recency effect on most measures. In the fast laboratory version there was neither primacy nor recency. That was a causal learning task rather than free recall, so it is not a failed replication of Murdock. It is a warning about how far the effect travels.

And structure changes things. Ata Karagoz and colleagues had 66 participants learn words grouped by hidden rules, so that shifts between rules created event boundaries, and then freely recall them [53]. Recall clustered by event, as predicted. But recall was worse for words encoded immediately after a boundary, which is the opposite of what theories of event cognition expected. Position within a structure matters more than raw position in a sequence. Geoff Ward and Lydia Tan reported something similar for meaningful lists. Using overt rehearsal across three experiments in 2023 with lists of up to 64 words, they found people rehearse a related earlier word when a new one arrives, sometimes reaching back over more than a dozen intervening items [54]. Once material has meaning, position stops being the only thing organising it.

So the honest answer is this. The curve is real outside the laboratory, and it is weaker, more conditional and easier to disrupt than the confident advice built on it suggests.

The Same U Shape, in a Machine

In 2024 a group of researchers led by Nelson Liu ran an experiment that is almost exactly the Murdock design, on a language model.

They gave the models a question and a stack of documents, exactly one of which contained the answer, and moved that document through the stack [55]. Nothing else changed. Accuracy traced a U, and the authors reached for the psychologists' vocabulary to describe it [55]. They call the two ends a primacy bias and a recency bias, and, as they put it, models "suffer degraded performance when forced to use information within the middle of its input context". GPT-3.5-Turbo's accuracy on the task can drop by more than 20 percent purely from where the answer sits, which is the paper's own figure.

The detail that makes it worth telling is what happened in the middle. With 20 and with 30 documents in the context, putting the answer in the middle dropped performance below the closed-book baseline of 56.1 percent, which is what the model scores when it is given no documents at all [55]. Handed the answer in the wrong position, the system did worse than it would have done with nothing.

GPT-3.5-Turbo accuracy with no documents and with only the answer documentNo documentsAnswer document only1009080706050403020100Accuracy percent

Those two bars are the floor and the ceiling from the same study. Bury the right document among nineteen others and the model falls under the lower bar.

Nobody built a primacy effect into these systems. It emerged from attention over a sequence. Nikolaus Salvatore and Qiong Zhang pushed that further in 2025, showing that sequence-to-sequence models with attention contain mechanisms mapping onto those specified in the Context Maintenance and Retrieval model of human memory [56]. Two systems built for different reasons, converging on one solution.

A machine showing a U-shaped curve does not tell you how your hippocampus works.

It does tell you the curve is not unique to biology, and that anything which processes a sequence and has to find something in it later may end up with the same weak spot.

What Follows, If You Are Trying to Learn Something

You will have met the advice: put your best point first and your second best last. Some of that survives scrutiny and some does not.

What does not survive is any advice that leans on the classic recency effect. That effect is gone after 30 seconds of counting [5]. If you revise something at the end of a session and then walk to the kitchen, the advantage you are counting on has already expired. Recency in the useful sense, the sort that lasts, is the long-term recency of the 1974 continuous distractor task, and that one depends on spacing rather than on being last [11], assuming Koppenaal and Glanzer are wrong that it is an artefact of the distractor.

What does survive is the observation that every session has a beginning and an end, and that more sessions means more beginnings. That is the same argument as the spacing effect, arrived at from a different direction, and it is the one piece of practical advice here with a large literature behind it.

A shorter list has less middle in it, which is the practical version of the same point. Murdock's boundaries put primacy over the first three or four items and recency over the last eight, so a list of 30 has about 18 items in the trough and a list of ten has almost none.

Breaking 30 into three tens does not shorten what you have to learn. It changes how much of it is ever in the trough. That is what chunking buys you in serial position terms.

The order you work through is not fixed either. Revise in the same sequence every time and the same items sit in the trough every time. Shuffle it and the trough moves, though nothing has tested that directly as a study technique, so treat it as an inference rather than a result.

The least comfortable point comes last. The curve is at its biggest when people have to produce the words rather than pick them out. Ask them to recognise the words instead and much of it flattens, though distinctiveness effects persisted even there in Neath's 1993 recognition study [17]. Test yourself by rereading and nodding and you are running the version of the task where position matters least and your sense of knowing is least reliable. That gap between recognition and recall is worth more to you than any positional trick.

There is also a real limit on how much of this is about the list at all. Karl Healey and Michael Kahana identified four factors in how people search memory, covering the tendency to start from the end, the tendency to start from the beginning, and transitions driven by time and by meaning, and those four together accounted for 83 percent of the variance in overall recall accuracy [57]. Two of those four are about position and two are not, which is a reminder that where an item sat is only part of what decides whether you get it back. Age changes the balance too. Michael Kahana and colleagues found in 2002 that the recency effect and the tendency to recall neighbouring items together come apart with age, so older and younger adults produce different curves for different reasons [58]. No single curve describes everyone.

What Is Settled and What Is Not

Some things you can say without hedging.

The U-shaped curve is real. It replicates across word lists, spoken lists and video lists, in animals and in patients, and something with the same shape turns up in language models reading long contexts.

A filled delay after a list abolishes the classic immediate recency effect, which two laboratories showed independently in the mid 1960s. Early items get more rehearsals than late items when rehearsal is spoken aloud and counted. And in group studies the shape of a curve has carried clinical information that the total score did not.

And the things you should not say without naming both sides.

Whether recency needs a short-term store at all is open, with Koppenaal and Glanzer on one side, Bjork, Greene and Nairne on the other, and Talmi's own work landing on both. Whether distinctiveness or retrieved context is the better replacement is open, and Murdock's 2008 objection to SIMPLE and the Farrell and Lewandowsky exchange with Howard are both worth reading before anyone declares it closed. Whether primacy is really about rehearsal or about the order people choose to recall in is open, with Lydia Tan and Geoff Ward's 2000 demonstration that a recency-based mechanism can produce primacy on the second side [59] and their later work on list length and output order behind it [60]. And whether the signatures used to identify a working memory store identify anything at all is now in question after the 2026 modelling work [27].

The middle deserves the last word. It looks like the boring part because it is the flat part. It is the only region taking interference from both sides, the only one where rehearsal has run out, and the only one a drug has been shown to damage while leaving both ends standing.

Which suggests the way to think about it is not that the middle is forgotten.

It is that the middle is the default, and both ends are exceptions. Being first is an exception. Being last is an exception. Being unusual is a third exception, which is why an isolated item escapes the trough.

Everything else is the middle. Most of what you learn today will be the middle of something.

Frequently Asked Questions

What is the serial position effect?

The serial position effect is the finding that your chance of recalling an item depends on where it sat in a sequence. Items near the start and the end are recalled better than items in between. Bennet Murdock mapped the standard shape in 1962 across 103 people and lists of 10 to 40 words, describing a steep advantage over the first three or four items, an S-shaped advantage over the last eight, and a flat stretch between them. The two ends are called the primacy effect and the recency effect. The flat stretch has no popular name, which is part of why it gets ignored.

What is the difference between the primacy effect and the recency effect?

Primacy is the advantage for items at the start of a sequence and recency is the advantage for items at the end. They behave differently under interference, which is what made them interesting. Counting aloud for 30 seconds after a 15-word list removes the recency advantage entirely and leaves primacy untouched, a result Glanzer and Cunitz published in 1966 with 46 men, and one Postman and Phillips had reported independently in 1965 with their own participants. Repeating items during the list helps early and middle positions and does little for the last few. The classic interpretation is that they reflect two different memory stores, and that interpretation is contested.

Why do we forget the middle of a list?

Three things go wrong at once. Rehearsal runs out, because by the middle of a list the set of items you are repeating is saturated and each new item gets one or two passes instead of the four or five the first item received, though whether that is the whole story for the primacy end is still argued. Interference arrives from both directions, since middle items compete with earlier items at retrieval and have their encoding context overwritten by later ones. And nothing about their timing makes them distinctive, because their neighbours sit at almost identical temporal distance. The middle is not simply the absence of what helps the ends. In 2025 biperiden, a drug used to model early Alzheimer's memory loss, was found to leave primacy and recency intact while impairing memory for the middle 10 words.

How long does the recency effect last?

The classic one is gone within about 30 seconds if that time is filled with a task. Ten seconds of counting removes most of it and 30 seconds removes all of it. There is a second kind that lasts far longer. If a distractor task follows every item in the list rather than only the last one, the advantage for recent items returns, a result Bjork and Whitten published in 1974. That version depends on the spacing between items relative to the delay before recall rather than on how recently the last item arrived. Whether it proves that no short-term store is involved is still disputed.

Does the serial position effect apply to AI language models?

It does, and the researchers who found it used the psychological terms. Liu and colleagues moved a single answer-bearing document through a language model's input context in 2024 and accuracy traced a U, which they described as a primacy bias and a recency bias. GPT-3.5-Turbo's accuracy can drop by more than 20 percent purely from the position of the relevant document. With 20 or 30 documents in the context, placing the answer in the middle dropped performance below the 56.1 percent the model scored with no documents at all. Nobody designed that in. It emerged from processing a sequence.