Introduction
Reading should not work. Writing was invented somewhere around 3,400 BCE, which is nothing on an evolutionary clock. There was no selection pressure for literacy, no ancestral advantage in decoding marks on clay, no time for a dedicated reading organ to grow. And yet almost every literate adult on the planet does it at roughly 250 words a minute, without effort, without conscious decoding, and without the ability to switch it off. Look at a word in a language you know and the meaning arrives before you decide to look.
Stanislas Dehaene and Laurent Cohen called this the reading paradox [1]. A skill too recent to be evolved is nonetheless implemented in nearly the same place in nearly every brain. In the year 2000, a team working in Paris put a name to that place: the visual word form area, a patch of left ventral occipitotemporal cortex that responds to written words far more than to almost anything else [2].
The name stuck. The agreement did not.
What follows is the story of that patch of cortex. It starts with a retired salesman in Paris who could write a letter and then could not read it back. It runs through a century of near silence, a naming, an immediate accusation that the name was a myth, and a run of studies since 2019 that have quietly changed the terms of the argument. Three research camps are still fighting over what this region actually computes, and this article does not pick a winner, because the field has not picked one either. What the field has done is converge on something more interesting than any of the three original positions.

The Salesman Who Could Write But Could Not Read
Joseph Jules Déjerine first met the man in 1887. In the neurological literature he is Monsieur C., a retired salesman, educated, articulate, and by all accounts entirely himself. One afternoon he was reading in his chair and the words stopped working. Not blurred. Not faint. Simply meaningless.
What made the case famous was not the loss. It was what survived it.
Monsieur C. could still speak fluently. He could still understand everything said to him. He could still write, spontaneously and to dictation, in a clear hand. Then he would look down at the sentence he had just produced and have no idea what it said. Déjerine published the case in 1892 under the heading of pure word blindness, cécité verbale pure [3]. Writing intact, reading destroyed. Two abilities that feel like one thing turned out to be two.
Here is where the textbook version of this story has drifted from the record. Generations of summaries claim Monsieur C. could still read numbers perfectly, giving a clean dissociation between letters and digits. When Daniel Bub, Martin Arguin and André Roch Lecours went back through Déjerine's original observations in 1993, they found the picture was messier [4]. Number reading was impaired too, just less severely than letter reading. The dissociation was a gradient, not a wall. It is a small correction with a large moral, which is that the founding case of reading neuroscience has been simplified in the retelling for over a century.
The autopsy told the rest. Damage to the left medial occipital cortex, damage to the posterior ventral occipitotemporal region, and critically, damage to the splenium of the corpus callosum. The angular gyrus, which Déjerine had already identified in an 1891 case as the seat of stored visual word images, was intact.
That combination gave him his explanation. Visual information from the left visual field reaches the right occipital cortex, then normally crosses to the left hemisphere through the splenium. Destroy the left occipital cortex and the direct route is gone. Destroy the splenium and the detour is gone too. The angular gyrus survives, undamaged and useless, cut off from every source of visual input. Déjerine had described a disconnection.
His 1891 patient had made the contrast possible. That earlier man had damage to the left angular gyrus itself, and he lost reading and writing together. Same family of symptoms, different lesion, different pattern. One case with the store destroyed, one case with the store isolated.
What Déjerine had produced, without the vocabulary for it, was a wiring diagram. He was not claiming that reading lives in one spot. He was claiming that reading depends on traffic between spots, and that you can destroy reading by cutting the road while leaving both destinations intact. That is a subtler idea than localisation, and it was probably too subtle for the period.
There is a second reason the case mattered and a third reason it was neglected. It mattered because it separated two things that feel welded together in ordinary experience. Writing and reading feel like one skill with a direction switch. Monsieur C. proved they are not. It was neglected because the anatomy could only be confirmed at autopsy, which meant every claim in the field had to wait for a patient to die, and each new case brought a lesion of a different size in a different place. Sample size in nineteenth-century neurology was one.
Then the idea went quiet for most of a century. Disconnection thinking fell out of fashion, and pure alexia became a curiosity in the back of textbooks rather than a research programme. It took Norman Geschwind's two-part treatise in Brain in 1965 to bring the framework back into serious neurology, arguing that complex deficits can arise from severing the white matter between intact regions rather than from destroying the regions themselves [5]. Déjerine's salesman had been right all along. Nobody had the tools to prove it.
A Name, and the Coordinates That Came With It
Functional imaging changed that. By the late 1990s researchers could watch the intact reading brain instead of reasoning backwards from lesions, and one region kept appearing.
The 2000 paper by Cohen, Dehaene, Naccache, Lehéricy, Dehaene-Lambertz, Hénaff and Michel in Brain is the one that named it [2]. The design was clever. They presented words to the left or right visual field in five healthy volunteers and, crucially, in two patients whose posterior corpus callosum had been surgically cut. In a split-brain patient, a word flashed to the left visual field has no route to the left hemisphere. If a left-hemisphere region still activates for words shown on either side in intact readers but only for one side in the split patients, you have located the convergence point of the two visual streams. That is roughly what they found, and they called the convergence point the visual word form area.
A follow-up in Brain in 2002 mapped its response properties in more detail and reported a peak that has been quoted ever since [6]. The 2003 review by Bruce McCandliss, Laurent Cohen and Stanislas Dehaene in Trends in Cognitive Sciences fixed the canonical value in the literature and framed the region as a case of visual expertise developing in the fusiform gyrus [7].
The coordinates are worth looking at directly, because their consistency is the strongest single argument the specialisation camp has.
Coordinates are in MNI space, in millimetres. Read the fourth row again. Those readers had never seen a letter in their lives [8]. We will come back to them, because that single row does more damage to the phrase "visual word form area" than any argument in the literature.
One qualification belongs here and applies to the whole table. These are group peaks. The region's exact position varies substantially between individuals, which is why group-averaged maps sometimes fail to show it at all and why individual-subject scanning has become the standard method rather than a refinement.

What the Region Actually Answers To
The response profile is one of the settled parts of this story. Present a fluent reader with real words and the region fires hard. Present pronounceable non-words that follow the spelling rules of their language and it fires almost as hard. Present strings of consonants that no language would permit and the response drops. Present false fonts, shapes with the visual complexity of letters but none of the structure, and it drops further.
Then there is the invariance, which is the genuinely strange part. The region treats radio and RADIO as the same object even though they share almost no pixels. Change the font, change the size, move the word to a different part of the retina, and the response holds steady. Dehaene, Naccache, Cohen and colleagues demonstrated case invariance in 2001 using masked priming, where a word flashed too briefly for conscious perception still primed its uppercase counterpart and still modulated activity in left fusiform cortex [9]. The reader never saw the prime. The region did.
That result is worth sitting with. A stimulus that never reaches awareness is nevertheless normalised across letter case by a piece of visual cortex, in a fraction of a second, automatically.
Sensitivity is graded rather than all-or-nothing. Binder, Medler, Westbury and colleagues showed in 2006 that activity in the left occipitotemporal region rises with bigram frequency, meaning how often the letter pairs in a string occur in the language [10]. The region is not asking whether something is a word. It is measuring how word-like the input is against a statistical model of the reader's own orthography.
Intracranial recordings back this up with a precision fMRI cannot reach. Lochy, Jacques, Maillard and colleagues placed electrodes directly in the left ventral occipitotemporal cortex of epilepsy patients in 2018 and recorded selective responses to letters and words at specific sites [11]. Not inferred from blood flow. Recorded from the tissue.
There is a useful way to think about why invariance is the hard part. A face-recognition system has to solve a similar problem, since a face seen from the side is a different image from the same face seen head-on, and the system has to decide they are one person. But faces come with a strong constraint. Two eyes above a nose above a mouth, in that order, always. Letters give the system almost nothing comparable. A lowercase letter and its uppercase partner can share no strokes at all, and the reader is expected to treat them as identical anyway. That mapping is arbitrary, learned, and specific to one culture's script. Whatever the region is doing, it has learned an equivalence that the visual world never taught it.
The timing matters too, and it is fast. Tarkiainen and colleagues used magnetoencephalography in 1999 to track responses to letter strings and found left inferior occipitotemporal sources active in the window between roughly 100 and 200 milliseconds after the word appears [12]. That is the same general window in which the visual system resolves complex patterns, a process explored in more depth in our piece on how fast the brain recognises visual patterns. Reading, in other words, is not a slow deliberate process that happens after seeing. It is folded into seeing.
The Letterbox Model
If the region is measuring word-likeness, how does it do it? Dehaene, Cohen, Sigman and Vinckier proposed an answer in 2005 that has organised a lot of subsequent work [13]. They called it the local combination detector model, and its logic is a hierarchy.
Neurons early in the ventral stream respond to oriented line segments. Slightly further along, populations respond to whole letter shapes regardless of font. Further still, to frequently co-occurring letter pairs. Further again, to common letter clusters and morphemes and finally to short whole words. Each stage has a larger receptive field and a more complex preferred pattern than the one before it. The system is building words the way object recognition builds faces, out of increasingly specific combinations of parts.
The prediction is spatial and testable. If the hierarchy runs from simple to complex, it should run from posterior to anterior along the fusiform.
Vinckier, Dehaene, Jobert, Dubus, Sigman and Cohen tested exactly that in Neuron in 2007 [14]. They built a graded stimulus ladder: false fonts at the bottom, then strings of infrequent letter pairs, then pseudowords, then strings of frequent letter clusters, then real words at the top. Posterior parts of the left ventral stream responded much the same to all of them. Move forward along the fusiform and the preference for word-like stimuli grew steadily stronger. The gradient was there.
This kind of hierarchical assembly, where the brain packages recurring elements into single larger units so that later stages have less to handle, is the same principle that makes chunking such an effective memory strategy. Reading is chunking implemented in cortex.
The diagram compresses a great deal. Real reading involves top-down influence at every stage, and the two routes are not cleanly separate. What it captures is the direction of travel through the ventral stream and the point at which a decision about word-likeness gets made.
Neuronal Recycling and Its Limits
So the brain has a word-processing hierarchy it could not have evolved. Where did it come from?
Dehaene and Cohen's answer in Neuron in 2007 is the neuronal recycling hypothesis, and the precise claim is narrower than the popular version [1]. They do not say the brain is a general-purpose learner that can be trained on anything. They say almost the opposite. A cultural invention can only be learned if it finds a pre-existing circuit that is already close enough in function to be useful and still plastic enough to be redirected. They call that a neuronal niche.
The consequences follow. If the niche is narrow, plasticity is bounded, and so is the possible variety of writing systems. Which leads to the corollary that gets least attention and is the most provocative: writing systems evolved to fit the brain, not the other way round. Scripts that suited the existing machinery survived. Scripts that did not were abandoned.
That sounds unfalsifiable until you look at what writing actually looks like.
One housekeeping note. Much of the vivid framing that circulates about neuronal recycling, the letterbox metaphor and the narrative sweep, comes from Dehaene's 2009 popular book rather than from the peer-reviewed papers [15]. The book is excellent and the theory is serious, but they are different kinds of source and should be cited differently.
Mark Changizi and Shinsuke Shimojo analysed 115 non-logographic writing systems across human history in 2005 and found that the average number of strokes per character sits at about three, and does not scale with the size of the character inventory [16]. A system with 20 characters and a system with 200 use the same stroke budget. In 2006, Changizi, Zhang, Ye and Shimojo went further, comparing the topology of letter shapes across roughly 100 writing systems, Chinese characters, and non-linguistic symbols against the distribution of contour junctions in natural scenes [17]. The junction types that letters are built from, the L shapes and T shapes and Y shapes, occur in written characters at close to the frequencies at which they occur in photographs of the world.
Marcin Szwed, Dehaene, Kleinschmidt and colleagues then showed in 2011 that the left occipitotemporal region is specifically sensitive to those line junctions, and preferentially so for words over objects [18]. Letters look like the natural world because the machinery that reads them was built to see the natural world. The brain's capacity for this kind of structural reorganisation, and its limits, is the subject of a longer treatment in our article on how the brain rebuilds itself.
The framework has critics, and they are not fringe. Michael Anderson's neural reuse proposal in Behavioral and Brain Sciences in 2010 argues that redeployment of neural circuits for new purposes is a general organising principle of the whole brain, not a special mechanism for recent cultural inventions [19]. If reuse is everywhere, then recycling is not explaining much that is specific to reading. Richard Menary's 2014 paper in Mind and Language pushes from a different direction, arguing that recycling underweights the role of learning-driven plasticity and cultural niche construction, and that the environment and the brain shape each other rather than the brain simply hosting a cultural guest [20].

The Myth Accusation
Three years after the naming, Cathy Price and Joseph Devlin published a paper in NeuroImage with a title that left no room for interpretation. "The myth of the visual word form area" [21].
Their case rests on two kinds of evidence, and both are worth stating precisely rather than in summary, because the argument is stronger than its reputation.
The first is neuropsychological. If the region is a dedicated visual word form processor, there should be patients whose damage is confined to it and whose deficits are confined to reading. Price and Devlin state that no such case has been reported. Real pure alexia comes with larger lesions and comes with company. The clean case that the theory predicts does not exist in the clinical record.
The second is functional, and it is a list. The same left mid-fusiform territory activates during tasks that involve no visual word forms whatsoever. Naming colours. Naming pictures. Repeating spoken words. Making manual responses to pictures of meaningless objects. And Braille reading in blind people, a finding Price herself co-authored with Büchel and Friston in Nature in 1998 [22]. If one region does all of that, then whatever it does is not specific to seeing words.
By 2011 they had built a positive alternative rather than just an objection. Their interactive account, published in Trends in Cognitive Sciences, treats ventral occipitotemporal cortex as an integration site working under predictive coding principles [23]. Perception combines bottom-up sensory evidence with top-down predictions generated automatically from prior experience. The region sits where visual features meet incoming predictions from phonological, semantic and motor systems. On this reading, orthographic specialisation is not built into the tissue. It emerges from the pattern of interactions, and pure alexia happens when the region is cut off from the predictions that normally reach it.
Alecia Vogel, Steven Petersen and Bradley Schlaggar pressed the empirical version of the complaint in 2014 in a paper titled "The VWFA: it's not just for words anymore" [24]. Their charge is methodological and it is sharp. Whether the region looks word-selective depends almost entirely on what you compare words against. Against a resting baseline or a flickering checkerboard, selectivity looks dramatic. Against other complex visual categories, it shrinks. They report conditions in which non-word and even non-letter stimuli drive the region as strongly as words do.
Supporting evidence arrived from other directions. Mei, Xue, Chen and colleagues found in 2010 that the region is involved in successful memory encoding of both words and faces, which is hard to square with a purely orthographic function [25].
The Paris group answered. Cohen and Dehaene's 2004 rebuttal in NeuroImage argues that the label was never a claim of exclusivity [26]. Nobody said the region does nothing but words. The claim is that it contains a reproducible population of neurons tuned to invariant orthographic regularities, and graded responses to other categories are exactly what expertise-driven specialisation in a general-purpose visual system should look like. Dehaene and Cohen restated the position more forcefully in 2011 [27]. McCandliss and colleagues had already conceded the basic point back in 2003, writing that object and face recognition also activate the region to varying degrees [7].
So the disagreement was never really about whether the region responds to non-words. Both sides agreed it does. The disagreement is about what that response means.
The Turn Nobody Predicted
Then the question changed shape.
In 2016, Zeynep Saygin, David Osher, Elizabeth Norton, Nancy Kanwisher, John Gabrieli and colleagues published a longitudinal study in Nature Neuroscience that reframed the whole problem [28]. They scanned children twice, once at age five before they could read and once at age eight after they could. At the first scan they measured white matter connectivity. At the second they looked for a visual word form area, and found one in 29 of 31 children.
The result was this. Where the region appeared in a given eight-year-old could be predicted from that same child's connectivity fingerprint at age five, three years before the child could read a word. Functional responses at age five predicted nothing. Connections did.
Nothing in the original debate anticipated that. Specialisation theorists had assumed the region emerges through experience. Interactive theorists had assumed it emerges through interaction. Both were partly right and both had missed that the address was already reserved.
Then the sample sizes grew. In 2019, Lang Chen, Demian Wassermann, Daniel Abrams, John Kochalka, Guillermo Gallardo-Diez and Vinod Menon analysed 313 adults from the Human Connectome Project and reported a double dissociation in structural connectivity [29]. Connections between the region and the lateral temporal language network predicted language ability but not visuospatial attention. Connections between the same region and the dorsal frontoparietal attention network predicted visuospatial attention but not language ability. Two distinct circuits, running through one patch of cortex, each predicting a different behaviour.
They named it a multiplex model, and coined a phrase that describes the whole post-2019 direction: connectivity-constrained cognition. The question stops being what the region is made of and becomes what it is wired to.
The relationship between attentional routing and what gets processed deeply is a recurring theme in cognitive neuroscience, explored further in our article on how attention gates memory.
The Precision Era, and an Unwelcome Surprise
Because the region's location varies between people, group averaging blurs it. The response was precision fMRI, which means scanning individual participants for long enough and across enough conditions to map their own region rather than an average one.
In 2024, Jin Li, Kelly Hiersche and Zeynep Saygin ran that method across four tasks and fourteen conditions and published the result in iScience [30]. The findings split the difference between the camps with unusual precision.
The critics were right that the region responds moderately to non-word visual stimuli. The proponents were right that it is nonetheless unique within ventral temporal cortex in the strength of its preference for written words. And then a third result that neither camp had built into its position: the region is the only category-selective visual area engaged by auditory language. It responds when a person hears speech, with no visual input at all.
Before that gets over-read, the same paper adds the qualifier. The language response is dwarfed by the visual response, and the authors are explicit that this does not make the region a core amodal language area. Their preferred description is a visual look-up dictionary for orthography.
There is also internal structure. Alex White, John Palmer, Geoffrey Boynton and Jason Yeatman showed in 2019 that word-selective cortex is not one uniform patch [31]. A more posterior portion behaves differently from a more anterior portion, with the anterior region acting as a processing bottleneck where parallel spatial channels converge. Yeatman and White's 2021 review in Annual Review of Vision Science pulls this literature together and treats reading explicitly as a meeting point of vision and language rather than as a property of either one [32].
That framing is worth pausing on, because it quietly dissolves the original question. For twenty years the argument was whether the region belongs to vision or to language. Treating it as the junction between them is not a compromise position. It is a claim that the question was badly posed, and that a region sitting at the interface between two systems will inevitably look like a poor example of either one when tested with methods designed to find pure cases.
It also explains a pattern that had been sitting in the literature the whole time. Both camps kept finding what they predicted, in the same tissue, using different baselines. That is exactly what you would expect if the region's job is to translate between two codes rather than to compute in one of them.
Reading Without Eyes
Now back to that fourth row in the coordinate table.
In 2011, Lior Reich, Marcin Szwed, Laurent Cohen and Amir Amedi scanned eight congenitally blind adults reading Braille words and comparing them against meaningless Braille patterns [8]. These are people who have never had visual experience of any kind. No letters, no shapes, no light.
The peak of activation landed at MNI -45, -58, -12. That is the visual word form area, within a few millimetres of where it sits in sighted readers, and the authors describe the anatomical consistency as astonishing.
Consider what has to be true for that result to happen. A region of visual cortex, in a person with no visual history, organises itself around reading delivered through the fingertips.
It was not a one-off. Ella Striem-Amit, Cohen, Dehaene and Amedi showed in 2012 that blind users of a sensory substitution device, which converts visual shapes into soundscapes, activate the same region when identifying letters through hearing [33]. Katarzyna Siuda-Krzywicka and colleagues then found in 2016 that even sighted adults trained on tactile Braille recruit the region for touch [34].
The word "visual" in the name is doing no work at all.
This is the point where all three camps end up moving in the same direction, which is why it matters more than any single experimental result in this article. The specialisation camp now describes a metamodal reading region whose location is fixed by its connections to language areas. The interactive camp reads it as confirmation that the region is an integration hub rather than a visual detector. The connectivity camp reads it as the cleanest possible demonstration of connectivity-constrained cognition. Nobody defends the strong visual account any more.
That said, the metamodal reading is contested in its strongest form. Judy Kim, Shipra Kanjlia, Lotfi Merabet and Marina Bedny argued in the Journal of Neuroscience in 2017, from evidence in blind Braille readers, that development of the region does in fact require visual experience, and that the blind pattern is not simply the sighted pattern arriving by another road [35]. Same population, opposite conclusion. This one is unresolved.

Same Address, Different Scripts
A related question. Does a Chinese reader use the same machinery as a French one?
For years the answer looked like no. Logographic scripts seemed to recruit different regions, and the finding was often reported as evidence that reading networks are culturally determined. Then Kimihiro Nakamura, Wen-Jui Kuo, Felipe Pegado, Laurent Cohen, Ovid Tzeng and Stanislas Dehaene ran a study published in PNAS in 2012 that controlled the confounds more carefully by using cursive handwritten stimuli [36].
Both Chinese and French readers engaged two systems. One recognises word shapes, which is the ventral occipitotemporal route already described. The other recognises handwriting gestures, running through left premotor and middle frontal regions. Reading by eye and reading by hand, and both groups used both.
Meta-analytic work supports convergence with real but modest variation. Donald Bolger, Charles Perfetti and Walter Schneider's 2005 analysis in Human Brain Mapping found a shared core network across scripts, with Chinese and Japanese kanji showing somewhat more bilateral activation and English showing more extended activation than the more transparent Italian orthography [37]. Universal structures plus writing system variation, as their title puts it.
The most recent word on this comes from higher field strength. Minye Zhan, Christophe Pallier, Aakash Agrawal, Stanislas Dehaene and Laurent Cohen used millimetre-scale 7-tesla imaging in 2023 to ask whether the region splits in bilingual readers of two different scripts [38]. Standard fMRI shows no separate script-selective regions. At higher resolution, multivariate methods can decode which script is being read from activity patterns in the same territory. The specialisation exists. It is just finer-grained than conventional scanning can see.
How the brain manages two orthographies at once connects to broader questions about how bilingual brains store their languages, and to the deeper puzzle of how the brain acquires language in the first place.
What Literacy Costs
Reading is not free. Something has to give up cortex.
The study that made this measurable is Dehaene, Pegado, Braga, Ventura, Nunes Filho, Jobert, Dehaene-Lambertz, Kolinsky, Morais and Cohen in Science in 2010 [39]. The design is what makes it powerful. They scanned three groups of adults in Portugal and Brazil: 10 who had never learned to read, 22 who learned as adults, and 31 who learned in childhood. Same task battery for all of them, including written words, spoken language, faces, houses, tools and checkerboards.
The gains were substantial. Literacy produced the expected left fusiform response to writing. It enhanced visual responses more broadly across fusiform and occipital cortex, reaching all the way back to primary visual cortex. It strengthened responses to spoken language in the planum temporale. And it opened a top-down route from written input to phonological processing.
Look at the reach of that list. Learning to read did not just add a word-recognition module. It changed how the brain responds to speech, a skill the participants already had, and it changed activity as far back as primary visual cortex, which is about as early in the processing chain as it is possible to go. A cultural skill acquired in adulthood reached down and altered the response properties of the brain's first visual relay.
The finding that undercuts the critical-period story is that most of these effects appeared in the adults who learned late. The brain that learns to read at forty is not doing something categorically different from the brain that learns at six.
Then the cost. In the same left ventral region, responses to faces were reduced, and face processing shifted rightward. Reading appeared to have taken territory from something that was already there. Dehaene-Lambertz, Monzalvo and Dehaene followed this longitudinally in PLOS Biology in 2018, tracking children through their first months of reading instruction and reporting that word responses emerge partly at the expense of face responses at that location [40]. Régine Kolinsky and Tânia Fernandes reported a behavioural counterpart in 2014, showing that learning to read interferes with identity processing of familiar objects [41].
This claim is contested, and it needs flagging clearly because it circulates in popular coverage as established fact.
Bruno Rossion and Aliette Lochy published a critical review in Brain Structure and Function arguing that the accumulated evidence gives little if any support to the idea that right-hemisphere face lateralisation is caused by competition with left-lateralised word recognition [42]. Their central objection is a timing problem that is hard to answer. Rightward face lateralisation is already present in infants and in illiterate adults, in people who have never competed with anything for that cortex. If the effect exists before the supposed cause, the cause is not doing the work. Christian Gerlach and Randi Starrfelt have argued from a different angle, reporting that face processing ability does not predict reading ability in developmental prosopagnosia [43].
Treat face competition as an open question with serious researchers on both sides, not as a settled finding.
What is not disputed is that literacy leaves structural traces. António Castro-Caldas, Karl Magnus Petersson, Alexandra Reis, Sharon Stone-Elander and Martin Ingvar's 1998 PET study in Brain compared 12 Portuguese women, six literate and six illiterate, and found that repeating pseudowords recruited a different network in the illiterate group, while repeating real words did not differ much [44]. A follow-up in 1999 reported that the posterior midbody of the corpus callosum was thinner in illiterate adults [45]. Manuel Carreiras and colleagues found grey and white matter differences between literate and formerly illiterate adults in Nature in 2009 [46]. Dehaene, Cohen, Morais and Kolinsky pulled the behavioural and cerebral consequences together in a 2015 review in Nature Reviews Neuroscience [47].
Learning to read reorganises the brain at a scale that shows up in a callosal cross-section. Small sample sizes throughout this literature, and worth remembering. Six women per group is six women per group.
The First Quarter Second
Reading is fast enough that its stages have to be measured in milliseconds.
The workhorse measure is an event-related potential component called the N170, or the N1 for print, a negative deflection over left occipitotemporal electrodes peaking somewhere around 150 to 200 milliseconds after a word appears. In skilled adult readers it is larger for letter strings than for symbol strings, and it is left-lateralised.
In children it looks different. Urs Maurer, Silvia Brem, Kerstin Bucher, Daniel Brandeis and colleagues reported in 2006 that the component is broader, later and more bilateral in young readers, and that coarse neural tuning for print peaks as children learn to read rather than being present beforehand [48]. Print tuning is built, not delivered.
How fast can it be built? Silvia Brem and colleagues gave pre-reading kindergarteners short training on letter-to-sound correspondences in 2010 and found that print sensitivity emerged in the occipitotemporal response as a result [49]. Not years of exposure. Learning which letter makes which sound was the trigger.
Intracranial work has clarified the sequence. Thomas Thesen, Carrie McDonald, Chad Carlson and colleagues recorded directly from the left fusiform gyrus in 2012 and described processing that is first sequential, moving from letters to words, and then interactive, with later stages feeding back [50]. Both camps can point at that result, which is part of why it is useful.
The approximate cascade runs from primary visual cortex at around 100 milliseconds, to occipitotemporal print tuning at 150 to 200, to lexical and semantic access at roughly 200 to 250. Those numbers are ranges rather than fixed values and they shift with age, task and stimulus.
The automatisation that produces those latencies is a specific kind of skill acquisition, one that follows the same general trajectory described in our piece on the psychology of expertise. Fluent reading frees processing capacity for comprehension, which is one reason the demands of decoding matter so much for what a reader can do with a text, a relationship covered in more detail in our discussion of cognitive load theory.
The Reading Circuit in Dyslexia
The clinical literature here is large, and one part of it is settled while another is genuinely open.
The settled part is that reduced activation in the left occipitotemporal region is a reliable finding in developmental dyslexia. Bennett Shaywitz, Sally Shaywitz, Kenneth Pugh and colleagues reported disruption of posterior reading systems in children with dyslexia in Biological Psychiatry in 2002 [51]. Eraldo Paulesu and colleagues had already shown in Science in 2001 that the pattern crosses languages, appearing in English, French and Italian readers despite very different orthographies [52]. Fabio Richlan, Martin Kronbichler and Heinz Wimmer confirmed the convergence through meta-analysis in 2009 and again in 2011 across both children and adults [53][54]. Maurer and colleagues found the electrophysiological version of the same thing in 2007, reporting impaired tuning of the fast occipitotemporal print response in dyslexic children [55].
The open part is what that reduction means, and it is a serious methodological problem rather than a detail.
A child with dyslexia has read far less text than a typically developing peer of the same age. Reduced activation in a region that is built by reading experience could therefore be a consequence of reduced experience rather than a cause of the difficulty. The standard tool for separating these is a reading-level-matched design, comparing dyslexic readers against younger children reading at the same level rather than against age peers, and some differences shrink under that comparison. The interactive account makes a specific prediction here too, since if occipitotemporal activation depends on top-down predictions from phonological systems, then weaker activation would be downstream of a phonological difficulty rather than a primary cause of it [23].
Evidence on the other side arrived recently. A 2025 preprint from Jamie Mitchell, Maya Yablonski, Hannah Stone, Jason Yeatman and colleagues reports that in dyslexic readers the word-selective region is more often small or absent, and, more strikingly, that this persists after intensive intervention [56]. Forty-four children with dyslexia completed 160 hours of reading intervention, with 46 additional participants as comparison, and were scanned repeatedly over the course of a year. Reading skill improved. Region size increased. The gap did not close.
Two things need saying about that study before anyone builds on it. It is a preprint, which means it has not completed peer review at the time of writing. And the intervention tested was a commercial reading programme, with the study conducted in collaboration with the company that owns it, which is a conflict of interest that readers are entitled to weigh. The finding is interesting and it is not yet independently replicated.
It is worth being precise about why this question is so hard to settle rather than treating it as a gap someone will eventually fill. The two hypotheses make nearly identical predictions about a cross-sectional scan. A child who reads poorly because their word-recognition region developed atypically, and a child whose word-recognition region is underdeveloped because they read very little, will both show reduced activation at age ten. Distinguishing them requires either catching the difference before reading instruction begins, or changing reading ability and seeing whether the brain measure follows. Both designs are expensive, slow, and rare, which is why a field with thousands of dyslexia imaging papers still has very few studies that speak directly to the causal question.
Nothing in this section supports any conclusion about any individual. Neuroimaging findings describe group averages, the overlap between groups is substantial, and reading difficulty is identified through behavioural assessment by qualified professionals rather than through brain scans.
When the Left Hemisphere Is Lost Early
One more result, and it is the cleanest natural experiment in the whole literature.
Anna Seydell-Greenwald, Natalya Vladyko, Catherine Chambers, William Gaillard, Barbara Landau and Elissa Newport published a study in the Journal of Neuroscience in 2025 that scanned 15 people who had suffered a left-hemisphere perinatal stroke, all of whom had developed right-lateralised language, alongside 14 healthy sibling controls [57].
In the patients, the word-selective region had developed in the right fusiform gyrus, mirroring its usual left-hemisphere position. Controls were left-lateralised as expected, with a group statistic of t(13) = 6.03 and p < 0.0001. Patients were right-lateralised, with t(14) = 3.35 and p = 0.002.
Here is the detail that makes it decisive. These were middle cerebral artery strokes, which spare the ventral occipitotemporal cortex fed by the posterior cerebral artery. The left-hemisphere tissue where the region normally forms was intact and functioning. It responded normally to places. It simply did not become the reading region.
The region followed the language system to the other hemisphere while its usual home sat available and unused. That is difficult to reconcile with an intrinsic left-hemisphere visual bias, and it is difficult to reconcile with the idea that the location is set by competition with face processing. It fits the connectivity account almost too neatly, which is itself a reason for caution, and the sample is 15 patients.
One Hundred and Thirty Four Years in Ten Lines
Ten entries for a story that has produced thousands of papers. What the compression makes visible is the rhythm: a long silence, a fast naming, an immediate objection, and then a slow reframing that came from methods nobody in the original argument had access to.
Where the Field Actually Stands
Three accounts are live. Laying them side by side is more honest than picking one.
No winner is declared here because none has been declared in the literature. What has happened since 2019 is a convergence rather than a victory. All three camps now agree that the region's identity is set substantially by what it is connected to, that it is not purely visual, and that individual variation is large enough to require individual-subject methods. The remaining disagreement is about what the region computes once you know what it is wired to, and that question is open.
It is also worth naming what changed methodologically, because the reframing was driven by tools rather than by arguments. The original dispute was fought with group-averaged fMRI, a method that blurs a region whose position varies from person to person and that makes the answer depend heavily on the comparison condition. Diffusion imaging let researchers ask about wiring rather than activation. Longitudinal designs let them ask what came first. Precision scanning let them map individuals instead of averages. Larger cohorts let them relate brain measures to behaviour with enough power to detect dissociations. None of those tools existed in usable form in 2003, which is part of why that debate could not be resolved on its own terms.
Four specific claims should be treated as contested rather than settled.
The first is the function of the region itself, disputed across all three camps as described above. The second is whether literacy displaces face processing, where Dehaene and colleagues argue yes and Rossion and Lochy argue that the timing evidence rules it out. The third is whether reduced occipitotemporal activation in dyslexia is a cause or a consequence of reduced reading, where reading-level-matched designs and the interactive account push toward consequence while the 2025 intervention preprint pushes toward a stable trait. The fourth is whether visual experience is required for the region to develop normally, where Kim, Kanjlia, Merabet and Bedny argue that it is and the Amedi and Cohen groups argue from the Braille data that it is not.
Against those four, several things are genuinely settled. The region exists and its location is reproducible. It is causally necessary for fluent reading. Print tuning appears in the 150 to 200 millisecond window. It is absent in pre-readers and emerges with instruction. And it activates for Braille and for spoken language, which means the word visual in its name is a historical accident rather than a description.
Conclusion
The visual word form area was named in 2000 for what researchers thought it was. Within three years the name was called a myth. Within eleven years, blind people who had never seen anything were activating it with their fingertips.
That sequence is not a failure of science. It is what the process looks like when it is working. A name is a hypothesis compressed into two or three words, and this one turned out to be wrong in an interesting direction. The region is real, its address is reproducible to within a few millimetres across languages and scripts and sensory modalities, and something in the brain reserves that address before a child can read a word. What sits there is not a letter detector. It is a piece of cortex positioned at the junction between seeing and language, doing a job whose description is still being written.
Monsieur C. sat in his chair in 1887 and watched the words stop working. He could still write. He simply could not read what his own hand had made. It took a hundred and thirty four years, functional imaging, split-brain patients, illiterate adults in Brazil, blind Braille readers, 313 connectomes and children scanned before they could read to build an account of what had happened to him, and the account is still unfinished.
Reading is the most recent thing the human brain does well. It is worth remembering how strange it is that it works at all.
Frequently Asked Questions
What is the visual word form area?
It is a region of the left ventral occipitotemporal cortex, near MNI coordinates -43, -54, -12, that responds more strongly to written words than to most other visual categories. It was named in 2000 by Cohen, Dehaene and colleagues and is sometimes informally called the brain's letterbox.
Is the visual word form area really visual?
Not entirely. Congenitally blind adults reading Braille by touch activate almost the same coordinates, and precision scanning in 2024 found that the region also responds to spoken language. Most researchers now treat the word visual in its name as a historical label rather than an accurate description of its function.
Did the human brain evolve to read?
No. Writing is roughly 5,400 years old, far too recent for a dedicated reading circuit to have evolved. The neuronal recycling hypothesis proposes that reading occupies pre-existing object-recognition circuitry that was close enough in function and plastic enough to be redirected toward written symbols.
What happens in the brain when a child learns to read?
Print sensitivity develops in left occipitotemporal cortex, appearing in the 150 to 200 millisecond window of the brain's response. Training studies show that learning letter-to-sound correspondences is enough to trigger it. The region is absent in pre-readers and forms as instruction proceeds.
Do scientists agree about what the visual word form area does?
No. Three accounts remain active: specialised orthographic expertise, general integration of vision with top-down prediction, and a connectivity-based view in which the region serves both language and attention circuits. Research since 2019 has moved the field toward connectivity, without settling what the region computes.




