Introduction
Picture a maths classroom in Singapore. The teacher hands out a problem about how spread out a set of numbers is, and nobody in the room has been taught how to measure that.
There is no formula on the board. There is no worked example. The students are put into small groups and told to invent something.
They do. Groups produce an average of about six different ways to capture the idea, some clever, some clumsy, all of them theirs. Not one group arrives at the standard deviation.
Then the teacher steps in, puts the students' attempts next to the canonical formula, and shows exactly where each one falls short. On the test that followed, those students beat the class that was simply taught the formula, and they beat them on the questions the lesson never covered [1].
That sequence has a name. Manu Kapur called it productive failure in 2008, and it is one of the more interesting claims in learning science, partly because it keeps almost working and partly because the fight over it has never resolved.
You have probably met a version of this idea already, worn smooth. Let them struggle. Failure is the best teacher. That framing misses the half of the design that does the work.
So this article gives you the effect size, because the number exists and almost nobody prints it, and the two randomised experiments that found the opposite result, with their numbers too. Then it points at something odd sitting between them: the best experimental case against productive failure and the best evidence for it predict the same result for the same children, and neither side cites the other on it.
What Productive Failure Actually Means
Strip away the slogan and the design is specific. It has two phases and both are compulsory.
In the first phase, students meet a problem targeting a concept they have not been taught. The problem is chosen so that their existing knowledge gets them somewhere and not all the way.
They generate and compare their own representations and solutions. They usually fail, in the narrow sense that they do not reach the canonical answer.
In the second phase, the teacher arrives and does something particular. Not a fresh lecture from scratch. A comparison. In his own 2015 account of the design, Kapur describes this phase as consolidation and knowledge assembly, in which the students' attempts are set beside the target concept feature by feature until the reason the canonical solution is built the way it is becomes visible [2].
Drop the second phase and you have not run productive failure. You have run a wasted lesson.
Here is what that looks like when a real course does it. The figure below is from a biology course at the University of British Columbia that swapped the order for two topics and left everything else alone.

Chowrira SG, Smith KM, Dubois PJ, Roll I. DIY productive failure: boosting performance in a large undergraduate biology course. NPJ Sci Learn. 2019 Mar 12; 4:1. https://doi.org/10.1038/s41539-019-0040-6. Fig. 1. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.
Look at what changes and what does not. Both routes carry the same three components: student activity, formative feedback, and an expert walkthrough.
The only thing that moves is where the walkthrough goes. In the productive failure section it comes last. In the active learning section it comes first.
That is the whole intervention. Not more time, not harder material, not smaller groups. The order.
It is usually filed with the desirable difficulties, the conditions that make learning feel worse while it is happening and leave you better off afterwards, though whether that is the right home for it is one of the arguments below. It is a cousin of the testing effect and of interleaving your practice instead of blocking it.
What makes it different from those two is that it does not change how much you practise or in what order the topics come. It changes when the explanation arrives.
Seventy-Five Students Who Got Everything Wrong
The founding paper came out in 2008 in Cognition and Instruction, and its own claim is smaller than the reputation that followed it. What Kapur claimed in 2008 was an existence proof: eleventh-grade science students, randomly assigned, working on Newtonian kinematics in a shared online space, and the finding was that unsupported struggle on complex problems could be productive at all [3]. Not that it beats everything. That it is not simply waste.
The first proper classroom test came a year later, and it is the one worth sitting with. Seventy-five students in a mainstream Singapore secondary school, all in seventh grade, spent two weeks on rate and speed [4]. Half worked in the ordinary lecture-and-practice cycle. Half solved complex problems in groups with no support at all until a consolidation lecture in the final lesson.
The productive failure half generated a rich variety of linked representations and methods. Of those seventy-five students, the ones left to struggle also failed, in groups and alone, and reported low confidence in what they had produced [4].
Sit with that for a second. The condition this whole literature is built on felt worse from the inside, at the time, to the people in it. That gap between how learning feels and what it produces is the recurring theme of this whole literature, and it is the same gap that makes re-reading your notes feel productive while teaching you almost nothing.
A 2012 study replaced the topic with the concept of variance and reported what happened on the test. Ninth-grade mathematics students in the productive failure condition generated many formulations and none of them was the canonical one, and they still outperformed the directly taught group on conceptual understanding and on transfer, without losing any procedural fluency [1].
That last clause matters more than it looks. Nothing was traded away. The formula was learned just as well.
The Condition That Should Have Won
Then Kapur ran the study that makes the whole thing strange.
He took 109 students, again seventh graders, again Singapore, all taught by the same teacher, and gave them three conditions instead of two [5]. Lecture and practice. Productive failure. And a third arm that was identical to productive failure except that students got instructional facilitation all the way through the problem-solving phase.
That third condition is the sensible one. It keeps the generative work and adds help. Anyone designing a curriculum reaches for it.
It lost. Across the 109 students, the productive failure group beat both of the others, on the well-structured problems and on the higher-order application problems, and produced greater flexibility in moving between representations [5]. This is not an isolated result. In two further quasi-experiments with 104 students and then 175 students, a guided version of the problem-solving phase, where students were steered toward the components of the canonical solution, did not improve on the unguided version [6]. Helping during the struggle is not a free upgrade. On the current evidence it is not an upgrade at all.
Think about what that rules out. The obvious way to make the method safer, by keeping the generative work and adding support to it, is the version that lost. Whatever the struggle is doing, your help interferes with it.
What The Failure Is Doing
So what happens in a student's head while they produce six wrong answers?
Kapur's account names four things. Your prior knowledge gets activated, and then differentiated, so you find out which of the ideas you already hold are relevant and which only looked it. Attention lands on the features that actually matter. It has to, because those are the features that keep breaking your attempts.
Those features then get explained and elaborated. In the consolidation phase they get assembled into the concept.
The best single account of the mechanism is a 2016 review that pulled the whole literature into one theory of when and how problem solving before instruction supports learning [7]. Its central point is that the exploration phase is not there to produce answers. It is there to make you notice a gap.
There is evidence for that in the process data. Across the two randomised studies Kapur reported in 2014, the number of distinct solutions a group generated predicted how much its members learned [8]. Not the quality of the solutions. The quantity of them. Generating more ways to be wrong meant learning more from the lesson that followed.
That fits a wider finding about errors. A 2017 review of the error literature concluded that errors committed with high confidence are corrected better than errors committed with low confidence, which is the opposite of what most teachers expect [9]. The effect has a name, hypercorrection, and a 2011 experiment found it survives a week-long delay, though the high-confidence errors do have a habit of coming back [10]. A 2018 classroom study found it holds outside the laboratory too [11].
Being wrong loudly is easier to fix than being wrong quietly.
The error literature also supplies a warning. A 2019 study found that generating a wrong answer before being shown the right one improves memory for the items themselves and not for the associations between them, which makes the benefit narrower than the enthusiasm around it [12]. This is a family of related effects, not one clean law.
The Number Nobody Prints
Search for productive failure and you will find page after page telling you it works, most of them gesturing at "a review of 53 studies" without ever saying what the review found.
Here is what it found. In 2021, Tanmay Sinha and Manu Kapur published a meta-analysis in Review of Educational Research covering 53 studies and 166 experimental comparisons, drawing on more than 12,000 participants by the authors' own account [13]. Problem solving before instruction beat instruction before problem solving with a Hedges' g of 0.36 and a 95 percent confidence interval running from 0.20 to 0.51. Take a moment on that number. It is moderate. It is not nothing and it is not a revolution. An effect of 0.36 means a typical student in the problem-solving-first group ends up ahead of about 64 percent of the comparison group [13].
Compare that against how the method is usually described. Most pages about productive failure call the evidence strong. Strong is not a number. Nought point three six is a number, and it is the one that belongs in the sentence.
When the design stuck closely to the principles of productive failure rather than merely putting a problem before a lecture, the effect ran from 0.37 to 0.58 [13]. Fidelity matters. Any problem-solving activity followed by any instruction is not the same thing. The same 2021 paper also estimated what the effect would be after correcting for publication bias and put it at 0.87, which is a much larger number and also a modelled estimate rather than a measurement [13]. It belongs in the record. It does not belong in a headline.
Better At What
Now the distinction that reorganises everything else.
Go back through the studies above and look at which outcomes moved. Procedural fluency, meaning your ability to carry out the taught procedure, comes out level. Two randomised controlled studies reported in 2014 found both orders produced high procedural knowledge, and the difference appeared only in conceptual understanding and in the ability to handle problems the lesson had never shown them [8].
The technical word for that second thing is transfer, and it is the outcome this entire method lives on.
The dissociation was documented before productive failure had a name, and again after. In a 2011 study, students who were told the concept first did better on problems like the ones they had practised and worse when the task required them to move the idea somewhere new, while students who invented first showed the reverse pattern [14]. Telling first buys fluency. Inventing first buys flexibility. You do not get both from one lesson.
Which gives you a way to predict when a teacher will conclude the method failed. If you evaluate it with a quiz that asks students to reproduce the procedure, you will find nothing, because there is nothing there to find. The gain lives in the questions you did not think to ask.
Five Hundred And Seventy-Four Students, And A Final Exam
Most of the productive failure evidence comes from lessons designed by learning scientists. That is a fair objection to it, and in 2019 a group at the University of British Columbia set out to test what happens when ordinary instructors run it themselves [15].
They took 574 students in a first-year cell biology course. Two sections, identical syllabi, identical exams. In one, 295 students got productive failure for two topics: DNA replication, and transcription and translation. The other 279 got that content taught the way the course already taught everything else.

Chowrira SG, Smith KM, Dubois PJ, Roll I. DIY productive failure: boosting performance in a large undergraduate biology course. NPJ Sci Learn. 2019 Mar 12; 4:1. https://doi.org/10.1038/s41539-019-0040-6. Fig. 3. Licensed CC BY, https://creativecommons.org/licenses/by/4.0/.
On the midterm covering those topics, the productive failure section scored 4.78 percentage points higher, with a 95 percent confidence interval from 2.19 to 7.36 and p < 0.001, an effect size of 0.32 [15]. The bars in the figure sit further apart than that, because 4.78 is the difference left after controlling for what the students brought in with them. Then the term went on. By the final exam the advantage you just read about had gone. The estimate dropped to 0.78 points at p = 0.668, which is nothing at all [15].
You will not find that sentence on most pages about this. It belongs there.
One result survives the same standard, and it points the opposite way to the usual worry. Lower-performing students in the productive failure section were 8.28 points ahead at the midterm at p below 0.001. On the final they were 6.83 points ahead at p = 0.060, which by the standard applied one paragraph up does not count [15]. The students you would expect to drown in an unsupported problem were the ones who gained.
The Two Experiments That Went The Other Way
Everything so far has been the case for. Now the case against, and it is serious, because it varies nothing except the order and randomises students to it.
In 2019, Greg Ashman, Slava Kalyuga and John Sweller published two experiments in Educational Psychology Review that tested the order directly [16]. The material was light energy efficiency. The participants were Year 5 students at an independent school in Victoria, Australia, approximately ten years old, none of whom had been taught conservation of energy. They were randomly assigned to receive explicit instruction first or problem solving first, and nothing else differed. In the first experiment, with 64 students, instruction first won on the questions resembling those used in teaching, t(62) = 2.25, p = .03, Cohen's d = .56. On the transfer questions the difference did not reach significance, t(62) = 1.89, p = .06, d = .47 [16].
The second experiment raised the element interactivity of the material deliberately, and used 71 students. Instruction first won on both measures this time: similar questions at t(69) = 2.41, p = .02, d = .57, and transfer questions at t(69) = 2.35, p = .02, d = .56 [16]. The authors report mean scores almost fifty percent higher for the instruction-first sequence.
The bars are almost level. Whatever was happening in those two classrooms, it was consistent, and it pointed the other way.
Their conclusion, in their own words, was that "for learning where element interactivity is high, explicit instruction should precede problem solving." That is a boundary claim, not a blanket dismissal, and it is worth reading carefully.
The Boundary Both Sides Are Describing
Here is the thing that stopped me while reading these two literatures next to each other.
The Ashman experiments used ten-year-olds. Year 5.
Now go back to the meta-analysis, the flagship evidence in favour, and read past the headline number to the moderators. Sinha and Kapur report in their 2021 analysis that grade level affected the result, and that for learners in second through fifth grade the effect sizes favoured instruction first [13]. The same reversal appeared for domain-general skills.
The strongest experimental case against productive failure was run on children in the age band where productive failure's own meta-analysis says it should lose. Those are not the same claim. The critics argue from element interactivity and the meta-analysis argues from grade level. But they predict the same outcome for the same children, and neither side reaches for the other when saying so.
It is worth pausing on how easily this goes unnoticed. Two groups publish in one field, cite each other, and never join the halves.
That is not a gotcha in either direction. A boundary that removes primary school is a large boundary, and the reversal was already sitting in the evidence for the method.
The research group most interested in that boundary went and tested it properly. Earlier work with young children, the authors point out, had tested a stripped-down version and then reported that the method does not work for children, so in 2019 that team ran 228 students in fifth grade through a design carrying both of the components those studies had skipped: contrasting the students' own solutions against the canonical one during instruction, and collaborating in small groups while problem solving [17]. Then they tested two hypotheses about it and found no empirical support for either. Problem solving first was not better for these ten-year-olds, and collaborating did not help them either.
So the most careful attempt to rescue the method for young learners did not rescue it.
The honest position, and the one you should carry away, is that the age boundary is real and its cause is unsettled. It might be that younger learners lack the problem-solving skill to extract anything from their own failures. The design-fidelity explanation looks weaker after that 2019 study, which had the fidelity and still found nothing.
Element Interactivity, Without The Jargon
The critics are not waving their hands. They have a theory, and it is worth understanding because it makes falsifiable predictions.
Cognitive load theory starts from a 1988 argument that conventional problem solving imposes a heavy load that has nothing to do with learning, because searching for a solution and building a schema are different activities competing for the same working memory [18]. From there came the expertise reversal effect, described in 2003: instructional support that clearly helps a novice can actively harm a more knowledgeable learner, because processing guidance they no longer need is itself a cost [19]. A 2025 meta-analysis found the effect holds up across the literature [20].
Element interactivity is the concept underneath both. It counts how many pieces of information a learner has to hold in mind at the same time because each piece only makes sense in relation to the others.
Learning the chemical symbol for sodium is low in element interactivity. You can hold it on its own. Balancing an equation is high, because every term you touch constrains every other term.
The 2016 argument from Ouhao Chen, Kalyuga and Sweller is that expertise reversal is a special case of element interactivity, and that instructional effects flip sign when interactivity crosses a threshold [21].
One study set out to test that rather than assume it, running the two sequences against each other with element interactivity manipulated under experimental control, first with 52 students aged around ten and then with 96 students aged around thirteen [22]. On material high in element interactivity the worked-example-first sequence won, and most clearly for the novices. On material low in element interactivity, with more knowledgeable learners, there was no difference between the two orders at all. The prediction held where it was made and nothing contradicted it elsewhere. A 2015 paper showed the variable is strong enough to suppress the testing effect, which is otherwise about as replicable as anything in this field [23].
So the prediction is clean. Raise the interactivity, or lower the learner's prior knowledge, and the advantage should move to instruction first. That is exactly what the second Ashman experiment set out to show, and it showed it.
You can see why this argument has legs. It does not require anyone to have faked a result or run a bad study. It says the answer genuinely changes depending on the material, which would explain why two careful research programmes keep reaching opposite conclusions.
A 2024 paper put the two frameworks head to head and asked whether difficulty moderates learning in the way the desirable difficulties account claims or the way cognitive load theory claims [24]. That comparison is still open. Anyone telling you it is settled is selling.
Where It Stops Working
Three boundaries have been mapped, and a fourth is arriving. None of them is small, and you will not find them on most pages about this.
Young learners are the first, and the evidence for that reversal has already been laid out above.
The second is domain-general skills. When the thing being taught is a general capability rather than a specific concept, the 2021 meta-analytic effect sizes favour instruction first [13]. The meta-analysis found no subset of that literature pointing the other way. If you are teaching a general skill, you are on your own.
That reversal is easy to skim past and hard to explain away. It is not a null result sitting in the middle. The advantage moves to the other condition, in the very meta-analysis usually cited as proof that the method works.
The third is everything outside science, technology, engineering and mathematics. In 2020, Valentina Nachtigall and colleagues ran the first serious test of that, using social science research methods as the content and two quasi-experiments with 212 students and then 152 students, all in tenth grade [25]. The productive failure students did not outperform the directly taught students. Not by a little. Not at all. A 2025 analysis of English oral learning pointed the same way and blamed the ill-structured material and interference from the learners' first language [26].
There is a fourth boundary that gets less attention and comes from inside the productive failure camp. The whole design assumes students arrive with enough prior knowledge to generate something. Two quasi-experiments on monohybrid inheritance published in 2017 found that assumption failing: students did not necessarily have the micro-level knowledge needed to generate representations for a concept that operates at several levels at once [27].
Give a student a problem they have no purchase on and you are not producing productive failure. You are producing the conditions in which failure teaches people that effort does not matter.
It Is Not Discovery Learning
Most teachers who know one thing about educational research know this one: minimal guidance does not work.
That belief traces to a 2006 paper by Paul Kirschner, John Sweller and Richard Clark, which argued that constructivist, discovery, problem-based, experiential and inquiry-based teaching all fail for the same reason, namely that novices lack the schemas to make sense of unguided exploration [28]. It drew an equally forceful reply in 2007, arguing that problem-based and inquiry learning are heavily scaffolded and therefore were never the minimally guided methods being attacked [29].
Productive failure is not what that paper attacked. The failure phase is always followed by explicit instruction. A method whose second half is a teacher-led explanation is not minimally guided.
The cleanest evidence on this point is a 2011 pair of meta-analyses covering 164 studies. Across 580 comparisons, explicit instruction beat unassisted discovery at d = -0.38 with a 95 percent confidence interval from -0.44 to -0.31. Across 360 comparisons, enhanced discovery, meaning discovery with feedback, worked examples, scaffolding or elicited explanations attached, beat the alternatives at d = 0.30 with an interval from 0.23 to 0.36 [30]. Same behaviour, opposite verdicts, and the difference is whether anything is attached to it.
So when someone tells you productive failure was refuted in 2006, the accurate answer is that a different method was. The real counter-evidence arrived thirteen years later, and it is the Ashman experiments.
Unproductive Failure
In 2016 Kapur published a framing that should be better known than it is, because it is the thing standing between the research and the slogan [31]. Failure and success are one axis. Productive and unproductive are the other. That gives four cells, not two.
Most classroom misapplication lands in the second row. A teacher hears that struggle is good, sets a hard problem, watches the class flounder, and then moves on to the next topic because time ran out. Every ingredient of productive failure was present except the one that makes it productive.
The awkward part is that the second row and the first row look identical from the back of the room. Same problem, same confusion, same twenty minutes of nobody getting it. What separates them happens after you stop watching.
The newest work has moved to scaffolding the instruction phase rather than the problem-solving phase, which is a direct bet on that diagnosis. A 2019 study had already shown the lever working: elaboration prompts during consolidation changed how much students took from their own errors, while the errors themselves stayed the same [32].
Katharina Loibl and Nikol Rummel put the same point negatively in 2015, describing productive failure as a defence against the double curse of incompetence, the situation where you do not know something and also do not know that you do not know it [33]. That connects directly to what happens when you start monitoring your own understanding instead of assuming it.
Do You Have To Fail Yourself
An obvious question follows from all this. If the point is to notice a gap, could you notice somebody else's gap instead and skip the unpleasant part?
Kapur tested exactly that question in his pair of randomised studies. In the second of them, published in 2014, students who watched and worked with the failed attempts of their peers outperformed the students who had been taught first, but they did not match the students who had failed themselves [8].
Other research groups are not convinced by that reading. A 2021 experiment set productive failure against two example conditions, one where students watched another student's whole problem-solving-and-failing process and one where they saw only that student's failed solution at the end [34]. Doing it yourself was not superior to either. The students who watched the complete process actually outperformed the students who did the work, and prior knowledge helped the watchers while doing nothing for the doers. The follow-up in 2022 compared first-hand attempts against watching somebody else's and found the two equally effective at preparing students for the instruction [35].
So one group has your own failure as necessary and another has it as replaceable, on overlapping designs. That is what an unsettled question looks like from the inside.
The stake is practical. If watching is enough, productive failure gets cheaper to run and easier to sell to a nervous teacher.
The Idea Is Older Than The Name
Productive failure did not appear from nowhere in 2008, and the history matters because it shows the idea being narrowed decade by decade into a specific two-phase sequence with known limits. Its immediate ancestor is a 1998 paper with a title that has aged well.
The 1998 paper argued that analysing contrasting cases before a lecture changes what the lecture can do, because the learner arrives with the right distinctions already noticed [36]. That is productive failure with the failure taken out and the preparation left in. An invention strand ran alongside it. A 2012 classroom study of 134 students added metacognitive scaffolding to the invention phase, aimed at the four things students reliably skip: exploratory analysis, peer interaction, self-explanation and evaluation [37].
Who It Helps Most
That worry deserves a direct answer, because it is usually the first objection raised. Surely the weaker students just sink?
The evidence says no, twice.
The UBC scale-up already gave one answer: the lower-performing students were the ones sitting 8.28 points ahead at the midterm [15]. The second answer is more interesting. In 2023, Kapur, Janan Saba and Ido Roll ran two quasi-experimental studies across two Singapore schools chosen because their students differed sharply in prior mathematics achievement [38]. The expectation was that the stronger school would generate richer solutions. It did not. Students who were very different in prior achievement turned out to be strikingly similar in inventive production, meaning the variety of solutions they could design, and it was inventive production rather than prior achievement that had the stronger association with what they learned.
Prior achievement predicts what you can already do. Inventive production predicts what you can still learn. Those are not the same measurement, and the second one turned out to be the more evenly distributed of the two.
That distinction is worth carrying out of this article even if you forget everything else in it. What a student can currently do and what a student can currently learn are different quantities, and schools are set up to measure the first one.
The design has since been carried well outside the maths classroom, with mixed and small results. In physics, a 2018 study reversed the usual undergraduate routine and found conceptual knowledge improved [39]. In sport, a 2023 study let learners work out their own initial practice before being coached and found it improved javelin throwing performance [40].
Read that spread of results honestly and it does not look like a method taking over education. It looks like a method finding its edges, which is what you want a young idea to do while people are still testing it.
Nobody Would Choose This
There is a last problem, and it is the practical one.
Students do not want this. In the first Singapore replication of seventy-five students, the ones who had been left to struggle reported low confidence in what they had produced, and they were right to, because what they had produced was wrong [4]. How a lesson feels is a poor guide to whether it worked. The wider literature has measured that blindness directly. A 2017 study found learners are metacognitively unaware of the benefit of generating errors, and so do not choose the condition that teaches them more when the choice is offered [41]. A 2020 survey of students and instructors turned up beliefs about learning from errors that do not match the evidence, with students far more likely than instructors to treat an error as a sign that the study session failed [42].
You can hear the shift happening in students' own words. A 2026 interview study followed 12 students through a mandatory computer science course in the first year of medical school [43]. One of them described the moment the strategy changed. "At first I felt like every error meant I'd failed. Then I realised the errors were the lesson. Getting through them was the actual learning." Another described what memorisation had left behind, saying they knew the syntax and still could not write anything that solved a problem.
That is the reframing productive failure is trying to engineer, and it is not automatic. Somebody has to tell you that the errors were the point, or you conclude the lesson was a waste.
One more study belongs here, because it ran where the stakes were real. Across two cohorts of about 88 students each, preparing for a high-stakes state algebra examination, one group took mini-tests whose errors then became the focus of teacher-guided feedback sessions while the other was taught the material explicitly throughout [44]. Learning ran faster per hour of teacher time in the errors condition. The benefit was inconsistent across the two cohorts, which the authors say plainly, and that is the right level of confidence to carry away from it.
Where It Wins And Where It Loses
Put the strongest studies in one place and the shape is clearer than any summary sentence.
These five were picked as the strongest studies rather than to make a point, and the pattern the meta-analysis reported still falls out of the "who" column. The wins are secondary and tertiary. The losses are primary, or outside STEM.
What Is Settled And What Is Not
Some of this is no longer in dispute.
The design has two phases and the second is not optional. Students in the exploration phase almost never reach the canonical solution, and that is the design working rather than failing. On procedural fluency the two orders come out level.
Unassisted discovery does not work, and productive failure is not unassisted discovery. Learners do not choose these conditions for themselves, and they feel worse while doing them. The effects, where they exist, are moderate.
The open questions are larger than the settled ones.
Whether the age reversal is about cognitive capacity or about study design has not been resolved. Whether generating your own failed solutions is necessary, or whether watching somebody else's would do, has evidence on both sides. Whether the desirable difficulties framework or cognitive load theory better explains the pattern is an active argument with no verdict. And whether the method transfers outside STEM has two negative results in classroom subjects against one positive result in motor learning, which is suggestive rather than conclusive.
The clearest signal against the enthusiasts is about guidance during the struggle. Two separate research groups looked for a benefit and neither found one.
Conclusion
Productive failure is a real effect with a moderate size, a set of hard boundaries, and a design that most people describing it leave out.
If you want the shortest accurate version: across the whole literature, putting the problem before the explanation buys about a third of a standard deviation in conceptual understanding and transfer, provided a teacher afterwards contrasts what the students produced against the right answer.
It does not help anyone execute procedures faster. It reverses for primary school children. Outside STEM it has failed twice in classroom subjects and worked once on a physical skill, which is suggestive rather than settled. And in the one large university course that tracked it to the end of term, the advantage was gone by the final exam.
What makes it worth your attention is not the effect size. It is what the effect size implies about how understanding gets built. You do not learn a concept by receiving it. You learn it by having somewhere to put it, and generating six wrong answers is one reliable way of building somewhere to put it.
The uncomfortable part is that this feels like failing, because it is failing. Nobody in that ninth-grade room had a good afternoon. The students produced wrong answers, knew they were wrong, and went home with nothing that worked. Then the lesson landed harder than it would have.
And one part of the argument is less live than it is written. The strongest experimental case against the method comes from two studies of ten-year-olds. The meta-analysis in its favour reports that the effect reverses for second to fifth graders. Those are different claims reached by different routes, and they make the same prediction about the same children, and neither side has said so. What is open is why. Most of the rest is still open, and anyone who tells you the age question is the crux of it has not read both sides.
Frequently Asked Questions
What is productive failure?
Productive failure is a two-phase teaching design. Students first try to solve a problem targeting a concept they have not been taught, generating their own solutions and usually failing to reach the right one. A teacher then consolidates the lesson by contrasting the students' attempts against the canonical solution. Both phases are required. The struggle without the consolidation is what Kapur calls unproductive failure, and it produces nothing.
Does productive failure actually work?
It works moderately, within limits. A 2021 meta-analysis of 53 studies and 166 comparisons found a Hedges' g of 0.36 in favour of problem solving before instruction, with a 95 percent confidence interval from 0.20 to 0.51, rising to between 0.37 and 0.58 when the design kept close to productive failure principles. The advantage appears on conceptual understanding and transfer, not on procedural fluency, so a quiz that only asks students to reproduce a procedure will find nothing.
When does productive failure not work?
Three reversals are documented. The meta-analytic effect sizes favour instruction first for learners in second through fifth grade and for domain-general skills. Two randomised experiments with 64 and then 71 ten-year-olds found instruction first won on questions like those used in teaching, and on transfer questions too once element interactivity was raised. Outside STEM, two quasi-experiments with 212 and 152 tenth graders found no advantage for productive failure in social science research methods.
Is productive failure the same as discovery learning?
No, and this is the most common confusion about it. Discovery learning leaves students to find the concept themselves. Productive failure always ends with explicit teacher-led instruction, and the whole design depends on that instruction happening. A 2011 meta-analysis found unassisted discovery does not help learners while enhanced discovery, meaning discovery with feedback and explanation attached, does. Productive failure belongs in the second category, which is why the well-known 2006 critique of minimally guided instruction was aimed at something else.
Doesn't letting students fail hurt weaker students most?
The evidence points the other way. In a study of 574 undergraduate biology students, the lower-performing students in the productive failure section were 8.28 points ahead at the midterm, at p below 0.001. A 2023 study across two Singapore schools with very different prior achievement profiles found that students generated similarly varied solutions regardless of prior achievement, and that the variety of solutions predicted learning better than prior achievement did.




