Introduction
Your professor just curved the exam. Did that help or hurt you?
Here is the short answer first. If your instructor took a hard test and added points to everybody's score, it helped you, probably a lot. If your instructor decided in advance that only the top 15 percent may receive an A, your grade no longer depends only on what you know. It depends on what the people around you know. You are not being measured. You are being ranked.
Both practices are called grading on a curve. They are opposites, and almost nothing written on the topic separates them, which is why so much of the advice you find is useless.
This article is about the second, because it is the one seventy years of research has something serious to say about, and because it reorganises a classroom around a rule most students never see written down.
The rule is this. If the number of A grades is fixed, your classmate's success costs you something. Not metaphorically. Arithmetically.
That arrangement has a name, and the man who named it in 1949 appears on none of the pages currently ranking for this topic. Once you have his term, the evidence stops looking like a pile of complaints and starts looking like a set of predictions about what happens when you take away someone's ability to be measured on their own.
Two of the most inconvenient findings in this literature cut against the argument this article makes. They are in here too, with the names attached.
Two Things Called a Curve and Only One of Them Is the Problem
Start with the confusion, because it makes half of what you read elsewhere wrong.
The first is additive. An instructor writes an exam that turns out too hard, the class average lands at 54, and the instructor adds points to every score or scales the top up to 100. Nobody loses. Your grade rises because the test was badly calibrated, not because a classmate stumbled.
The second is distributional. Before anyone sits the exam, a policy fixes what proportion of the class may receive each letter. The class average is now irrelevant and only your position matters. If everyone improves, the same number of A grades are handed out. If you improve and three classmates improve more, your grade can fall while your knowledge rises.
Only the second creates the structure this article is about. When researchers write about norm-referenced grading, this is what they mean.
You can tell them apart with one question. Can the whole class get an A? If yes, it is a boost. If no, it is a quota. Hold on to that question. It is the most useful thing in this article.
What Festinger Actually Said
Leon Festinger published a theory of social comparison processes in 1954 [1]. It is one of the most cited papers in social psychology, and it is usually summarised in six words: people compare themselves to other people.
That summary is not wrong. It is the wrong half.
Festinger's first hypothesis is that humans have a drive to evaluate their own opinions and abilities. His second is the one that matters here, and it says nearly the opposite of the popular summary. When an objective non-social standard is available, people use it. They compare themselves to other people only to the extent that such a standard is missing.
Read that again with a gradebook in your hand. A criterion-referenced grade is an objective standard. It says you answered 83 percent of the questions on cellular respiration correctly, and that number does not move when a classmate has a good week. A norm-referenced grade is not a standard at all. It carries no information about what you know, only your position in the room. Robert Glaser drew that distinction in 1963 and named the two approaches [2], arguing that a relative score tells you who is ahead and nothing about whether anyone has learned anything.
So the curve does not make students compare themselves. They were already doing that. Pieternel Dijkstra and colleagues reviewed classroom comparison in 2008 and found it close to continuous, because school is the one place where dozens of people perform the same task under the same conditions all day [3]. Their review also found students prefer to compare upward, and that upward comparison tends to raise performance while lowering how capable you feel.
Pascal Huguet and colleagues showed the preference directly in 2001 [4]. Given a free choice of whose work to look at, students picked a classmate performing slightly above them, and their own subsequent performance improved.
The comparison instinct is not the villain here. Left alone, it pulls students upward.
What the curve does is remove the alternative. It takes away the objective standard Festinger said people would rather use, leaves the peer distribution as the only yardstick, and then makes that yardstick official.
The Man Who Defined What a Curve Is
Five years earlier, in 1949, Morton Deutsch published two papers that between them defined the structure a quota grading policy creates [5] [6].
His definition is arithmetic rather than emotional. A situation is competitive when the participants' goals are negatively correlated, so one person reaches their goal only if others fail to, and cooperative when goals are positively correlated. He called this goal interdependence and predicted what each arrangement produces before anyone measured it in a school. Negative interdependence, he argued, produces oppositional interaction, withholding of information, and distrust between people with no other reason to distrust each other.
A fixed grade quota is negative goal interdependence written into a syllabus. That is not an analogy. It is the definition, applied.
You can read a dozen articles about grading curves and a dozen more about social comparison and never meet the idea that explains why the first produces the second.
Deutsch also ran the experiment. His second 1949 paper put student groups into cooperative or competitive reward structures, and the cooperative groups communicated more, coordinated better and accepted each other's ideas more readily.
Seventy years of work followed. David and Roger Johnson counted more than 1,200 studies in the social interdependence tradition by 2009 [7], which is an unusually large evidence base for an idea in educational psychology and means the claims below do not rest on one clever experiment.
What Happens to the Room
The first place a quota shows up is whether students help each other.
In 1981 Johnson and colleagues published a meta-analysis of 122 studies containing 286 findings on cooperative, competitive and individualistic goal structures [8]. Cooperation outperformed interpersonal competition on achievement and also outperformed working alone, while competition measured against working alone was not reliably better at all.
That last result is the one people find hardest to believe. It does not say competition makes you worse. It says that once you compare competing against a peer with simply working alone, the advantage people assume competition brings did not appear.
The careful modern version is Cary Roseth, David Johnson and Roger Johnson's 2008 meta-analysis, which pooled 148 independent studies covering 17,000 adolescents across 11 countries [9]. Cooperative goal structures beat competitive ones on achievement at d = 0.46, rising to d = 0.57 when the weaker studies were excluded. On positive peer relationships the difference was d = 0.48. Competition against individualistic effort came out at d = 0.20 on achievement and d = 0.03 on peer relationships, neither statistically significant.
An effect size near d = 0.5 in education is not a rounding error. It is roughly what a well-designed intervention produces, and here it comes from nothing but the rule about who can win.
There is a catch, and it is what schools get wrong. Robert Slavin's reviews found methods combining group goals with individual accountability produced a median achievement effect near +0.32, more than four times the median near +0.07 for group work without those two features [10]. Seating four students at one table is not the intervention. Making the group's outcome depend on each member's own learning is.
The Study Everybody Cites and What It Actually Measured
The claim you will meet most often is that curved grading makes students stop helping each other. It is the emotional centre of nearly every argument against the practice.
The best modern evidence for it is a preregistered experiment by Jingxuan Liu, Michelle Wong and Bridgette Hard, published in 2024 with 547 students [11]. Participants read a description of a course graded either norm-referenced or criterion-referenced, then reported what they expected. The norm-referenced version produced higher performance-goal orientation, lower mastery orientation, lower expected self-efficacy, and fewer expected help-giving and help-seeking behaviours. Help-giving moved the most.
Now the qualification, from the authors themselves. The paper states it was unable to examine the impact of grading policies on actual student achievement. Participants reacted to a syllabus for a course they will never take, so what was measured is what students expect a curve to do to them, not what a curve did.
Expectations matter, because a student who expects a hostile room behaves differently on day one. But it is not the same as watching help disappear. The behavioural evidence for cooperation collapsing is older and indirect, resting on Deutsch and the Johnson meta-analyses.
Ruth Butler's 1988 experiment is the sharpest thing in this literature and it is nearly forty years old [12]. Twelve classes of fifth and sixth graders were randomly assigned written comments, numerical grades, or both, with interest and performance measured for 132 children across three sessions.
Comments alone produced the highest interest and performance, on a convergent task and a divergent one, among high and low achievers alike. Grades alone did worse. Grades with comments attached did about as badly as grades alone.
The feedback did not stop being useful. Adding a normative score next to it stopped the student from using it. Once a number told them where they stood, the sentences telling them how to improve became decoration.
The Pond You Happen To Be Swimming In
The second thing a curve does is change how good you think you are.
In 1984 Herbert Marsh and John Parker asked a question with an unusually memorable title: is it better to be a relatively large fish in a small pond even if you do not learn to swim as well [13]. Students of equal measured ability reported lower academic self-concept when surrounded by higher-achieving classmates. The big-fish-little-pond effect has been one of the most replicated findings in educational psychology ever since.
The size is worth stating carefully, because the public conversation about grading curves contains no effect sizes at all.
Marsh and Kit-Tai Hau tested it across 26 countries using PISA 2000 data, roughly 4,000 students per country [14]. The effect was negative in every single one of the 26 education systems, with a mean standardised coefficient of about -0.20 and a standard deviation of 0.08 across countries. It held at all ability levels.
A 2018 meta-analysis by Junyan Fang and colleagues pooled 33 studies covering 1,276,838 students [15]. The pooled coefficient was -0.28, with a 95 percent confidence interval from -0.32 to -0.24. It was stronger in high school at -0.32 than in college at -0.23, and it varied by region, running about -0.35 in Asia, -0.30 in Europe and -0.20 in North America.
More than a million students, negative everywhere anyone has looked. That is about as settled as educational psychology gets.
The bars show absolute values, so both big-fish-little-pond coefficients are negative in the original papers.
Two recent refinements make the finding more relevant to grading rather than less.
The first is local dominance. Marsh and colleagues showed in 2014 that the comparison doing the work is the nearest one, with the classroom frame mattering more than the school frame [16], and Malte Jansen and colleagues sharpened that in 2022, finding the referent is a subset of nearby peers rather than the class average [17]. Which is why a curve applied section by section bites harder than any school-wide policy. It publishes the local comparison, the one already doing the damage.
The second is experimental. Almost all of this literature is observational, leaving room for the objection that selective schools differ in a hundred ways. In 2023 Lisa Hasenbein and colleagues put 353 students into an immersive virtual reality classroom and varied how often the virtual classmates raised their hands [18]. Eye-tracking showed the students actively watching their peers' achievement behaviour rather than merely sitting near it. Pascal Huguet and colleagues had already shown in 2009 that measured comparison level statistically mediates the effect [19], so the damage runs through comparison itself rather than some other property of demanding schools.
Rank Is Not a Feeling
Educational psychologists measure self-concept, a self-report. Economists asked whether ordinal rank changes what happens to you afterwards, using administrative data on whole national cohorts.
Richard Murphy and Felix Weinhardt published the landmark study in 2020 [20], using 2,300,000 English students across 14,500 primary schools and 6,800,000 student-subject observations. Because children study several subjects, they could compare the same child's rank in one subject against their rank in another while holding ability constant.
A one standard deviation increase in primary school rank raised age-14 test scores by 0.084 standard deviations, about 2.35 national percentiles. Bottom of the class to top was worth almost 7.9 percentiles. Confidence rose by about 0.2 of a standard deviation on a five-point self-rating.
The authors put the size in terms a school leader can use. The rank effect is comparable to a year with a teacher one standard deviation above average, and about equal to the harm of four disruptive classmates. It lasted too, predicting which subjects students continued at A-Level seven years later. Benjamin Elsner and Ingo Isphording found the same shape in another national cohort, where higher ability rank raised students' expectations, effort and completed education independently of measured ability [21].
Read that next to a fixed grade quota and the implication is uncomfortable. A curve does not merely describe rank. It manufactures rank, sharpens it, and hands it to the student in writing.
One complication stops this being simple, and it comes from the same research team twice. Ghazala Azmat and Nagore Iriberri found in 2010 that telling high school students how they stood relative to the class average raised performance by roughly 5 percent [22]. When the same team studied college students over three years, repeated rank feedback reduced the number of exams passed, with the decline concentrated among students who had overestimated how well they were doing.
Short-term rank information can motivate. Sustained rank information appears to do something else.
The Two Findings That Should Slow You Down
An article that reported only the evidence above would be misleading, and every page currently ranking makes that mistake. Two results cut hard against the argument, and both are among the most rigorous work in the field.
The first is Kou Murayama and Andrew Elliot's 2012 meta-analysis in Psychological Bulletin [23]. They combined two meta-analyses with three primary studies and asked directly whether competition improves or damages performance.
The answer was neither. There was no net relation.
The explanation is the interesting part. Competition pushes students toward performance-approach goals, associated with better performance, and performance-avoidance goals, associated with worse. The two run in opposite directions and cancel. So the question most people ask, does competition help learning, is malformed. It depends which of the two responses a given student has.
David Johnson, Roger Johnson and Cary Roseth published a rebuttal in the same issue arguing the two mediators cannot be cleanly separated and the control condition was never well defined, and that their own 148-study analysis had already found cooperation ahead at d = 0.46 [24]. This is the cleanest named disagreement in the field, both sides published side by side, and it appears on none of the popular pages.
The second result is more direct. Moritz Fleischmann and colleagues used a Swedish natural experiment in which some municipalities graded students and others did not, with 9,104 students [25]. If the curve manufactures comparison, the big-fish-little-pond effect should be weaker where nobody hands out grades.
It was not. The effect was the same in graded and non-graded students.
Students compare whether or not an institution issues a ranking. Removing the grades removed the receipt, not the comparison.
Alli Klapp followed 8,558 students in Sweden and found the same pattern with a twist [26]. Being graded in Grade 6 had no main effect on Grade 7 achievement. But the harm was not absent, it was concentrated, falling on students with lower cognitive ability and on boys.
That pattern recurs, and it is the most useful thing to take from this literature. These policies rarely move the average. They move who bears the cost.
What It Sounds Like From Inside
Everything above is measured from the outside. Here is how students describe it themselves.
Edward Finch surveyed medical students at the University of Sheffield and published in 2026 [27]. 85 of 1,384 students responded, a rate of 6.1 percent, and 38 students left written comments. That is a small self-selected sample, so treat the percentages lightly and the sentences seriously.
One student explained their motivation in a line achievement goal researchers have spent thirty years building questionnaires to capture.
"A lot of my motivation comes from not wanting to be ranked last or in the bottom 50."
That is performance-avoidance, described by someone living inside it rather than scored on a scale. It is the exact mechanism Murayama and Elliot identified as the reason competition's benefits cancel.
Another student had watched Deutsch's prediction happen.
"Competing as an individual, can induce more stress in a person and I have seen people become slightly more hostile towards their peers, so they can show off how much better they are."
A third called competition fine in anonymous low-stakes quizzes and harmful when fixed on a number such as an end-of-year ranking. Students can tell a comparison from a quota. It is mostly the writing about grading that cannot.
One more result deserves a moment. Students rated themselves as less competitive than their peers. Everyone believes they are the calm one in a room of competitors.
Does a Curve Make Students Cheat
This accusation gets made most often and supported least well, so it needs care.
The prevalence figures are not in doubt. Baldwin and colleagues surveyed second-year students at 31 American medical schools, with 2,459 of 3,975 students responding [28]. 39 percent had witnessed cheating in their first two years, while 4.7 percent admitted cheating in medical school against 40.5 percent who admitted it in high school. The best single predictor was having cheated before. Sunčana Kukolja Taradi and colleagues found starker figures in Croatia in 2010 [29]. Of 761 students, 508 students responded, more than 99 percent reported at least one form of academic dishonesty and 78 percent admitted frequent cheating. Three had ever reported a classmate.
Those numbers describe how common cheating is. They do not show a curve caused it.
The link to classroom structure comes from elsewhere. Eric Anderman and colleagues reported in 1998 that early adolescents who admitted cheating were more likely to see their classrooms as focused on performance and extrinsic rewards rather than learning [30]. That relationship is correlational and the paper presents it as such. Caroline Pulfrey and colleagues added part of the mechanism in 2011, showing grades push students toward performance-avoidance goals through a loss of autonomous motivation [31], the same route Murayama and Elliot found cancelling competition's benefits.
So the honest version is narrower than the one you usually read. Classrooms that make relative standing the point are associated with more cheating, and a fixed quota is a strong way to do that. Nobody has run the experiment that closes the gap.
Where Curves Actually Live
Most writing treats the curve as one hard professor's quirk. It is closer to an institution.
American legal education is the clearest case. A survey of 186 law schools published in 2025 found 65 percent require grade normalization in first-year courses, with a further 11 percent advising it. In 1975 the figure was about 3 percent. That is not a teaching preference. It is a system, and it grew inside one professional lifetime.
Medicine ran the experiment in the other direction, which makes it the most informative case.
Darcy Reed and colleagues surveyed 2,056 students across seven United States medical schools in 2011, with 1,192 students responding [32]. Students at schools using three or more grade categories had a burnout odds ratio of 2.17, with a 95 percent confidence interval from 1.41 to 3.35, compared with students at pass and fail schools. The odds of having seriously considered dropping out were 2.24, from 1.54 to 3.27. Stress and emotional exhaustion scores were higher by similar margins.
The detail that makes this persuasive is what did not predict wellbeing. There was no association with the number of exams, the didactic or clinical time, or the length of vacation. The grading scheme predicted it. The workload did not.
Daniel Rohe and colleagues found the same pattern at Mayo in 2006, with lower perceived stress and higher group cohesion under pass and fail, and no difference in national licensing exam performance [33]. Laura Spring and colleagues reviewed the evidence systematically in 2011 [34]. All four studies examining wellbeing found improvement, five found no significant difference in objective academic outcomes, and two found residency program directors regarded pass and fail transcripts as a disadvantage when selecting candidates.
That last finding is the beginning of the real problem.
A counter-case deserves printing. Susan McDuff and colleagues found exam performance from year one to year two fell by 3.25 percentage points in a pass and fail class at San Diego, against 1.65 points under constant grading [35], while licensing scores were unaffected at 231.72 against 229.95. Removing the tiers may cost some measured effort while leaving the outcome that matters unchanged, and Robert Bloodgood and colleagues reported a similar mixed picture in 2009 [36].
What Happens When You Delete a Ranking
In January 2022 the United States Medical Licensing Examination stopped reporting a numerical score for Step 1 and began reporting pass or fail, the largest deliberate deletion of an educational ranking in recent memory.
The ranking did not disappear. It moved.
Matthew Pontell and colleagues surveyed surgery program directors, and with the Step 1 number gone a majority said they would lean harder on whatever rankable signals remained, including the still-scored Step 2 examination and the applicant's school [37]. Later work in orthopedics found the same displacement [38], and Eric Warm and colleagues named it in 2024, calling it the shadow economy of effort [39]. Remove the official competition and an unofficial one grows in its place, less visible and less fair, because the informal signals favour students who already knew how the system works.
That is the most practical lesson in the whole literature and it generalises well beyond medicine. Selection pressure migrates to whatever distribution is still visible. If a scarce prize sits at the end, abolishing one ranking mostly decides which ranking replaces it.
The Best Argument for the Curve
An article this long attacking a practice owes the reader its strongest defence, stated properly rather than knocked down.
If an instructor must grade against fixed absolute cutoffs, they are pushed to write exams most students can mostly complete, because an exam producing a class average of 48 would fail everyone under an absolute scale. Easy exams do not discriminate at the top. A curve frees the instructor to write genuinely hard questions, because performance is read relative to the cohort rather than against a number chosen in advance.
That argument is not silly and it is answered nowhere on the first page of search results. Instructors who curve are often doing it precisely because they refuse to dumb down the assessment.
The answer is that the calibration problem is real and the curve is a poor solution to it. Even instructors who map every question to a learning goal find their exam averages landing where they did not intend.
But you can keep hard exams without a quota, and the method is boring and effective. Publish an absolute scale at the start of the course, then reserve in writing the right to move the cutoffs down but never up. A student who earns 85 percent knows on day one that this is an A. If the exam turns out brutal, the cutoff falls to 78 and everyone benefits. Nobody's grade ever depends on a classmate.
That keeps the discrimination the instructor wanted and removes the negative interdependence. It is also the additive curve from the opening section, promised in advance instead of granted as a favour.
How a Curve Destroys Its Own Measuring Stick
One final argument against norm-referenced grading has nothing to do with wellbeing, and may be the strongest. Ranking systems decay.
In 1987 John Jacob Cannell surveyed standardised achievement testing across the United States and reported something arithmetically impossible [40]. All fifty states reported their elementary students above the national average on nationally normed tests. Not most. All of them.
The norms were stale, the tests were reused, and schools had every incentive to teach to them. The yardstick had not measured anything for years and nobody noticed, because everyone using it liked the answer.
That is what happens to any norm-referenced instrument once the measured can influence it. College grading has done the same more slowly. Stuart Rojstaczer and Christopher Healy assembled grade data from 135 institutions and found A grades had reached 43 percent of all letter grades awarded, about 28 percentage points above 1960, when C was the most common grade in American higher education [41]. A 2026 study of one research-intensive university tracked 22 years of graduate grading and found the same climb [42].
Jeffrey Denning and colleagues showed what inflation does to the signal [43]. Across two national cohorts the eight-year graduation rate rose 3.77 percentage points while first-year grade point average rose from 2.44 to 2.65 and precollegiate mathematics percentile fell from 58.88 to 55.93. Everything observable except grades predicted graduation should have declined by 1.92 points. First-year GPA alone explained 95 percent of the change. Grades also rose for students taking identical unchanged final examinations in required science courses.
Grades rose. Preparation fell. Completion followed the grades.
The earliest entries come from a 2025 law review survey of grading history rather than the 1920s papers themselves, and are reported on that authority.
Benjamin Bloom made the argument from first principles in 1968, and it remains the sharpest thing written on the subject [44]. The normal curve is the distribution you expect from chance and random activity. Education is purposeful. If teaching works, results should not be normally distributed, so a normal distribution of grades is evidence that the teaching failed.
Attempts to force the distribution back down have their own record. Princeton set a target in 2004 that no more than 35 percent of grades should be in the A range, and the A share fell from 46.0 percent to 40.9 percent in the first year. A faculty committee found no measurable damage to graduate admission, fellowships or employment. The faculty abolished the policy anyway in October 2014.
Wellesley College's experience should give any administrator pause. Kristin Butcher, Patrick McEwan and Akila Weerapana studied a cap on average grades and found enrolments and majors fell, student ratings of professors fell, and the racial gap in grades widened [45].
A policy introduced in the name of rigour made the outcome less equitable. Worth remembering before treating any grading reform as virtuous.
Who Self-Selects Into a Tournament
A quota does not only measure people. It decides who puts themselves forward.
Muriel Niederle and Lise Vesterlund ran an experiment in 2007 with 80 participants who performed a task under a piece rate and then chose whether to be paid by tournament [46]. 73 percent of the men chose the tournament. 35 percent of the women did. There was no gender difference in performance.
Part of what a competitive scheme selects for is willingness to compete rather than ability to perform, and any grading policy that turns a course into a tournament inherits that.
The pattern recurs. Klapp's Swedish cohort found the damage concentrated on boys and lower-ability students, Wellesley's cap widened racial gaps, and Liu and colleagues went looking for a clean equity story in 2024 and found their hypothesis inverted [11]. The equity picture is mixed, and anyone telling you it points cleanly one way has not read it.
The Same Machine, Rebuilt in Software
If you think this is a problem about old gradebooks, look at what replaced them. A leaderboard performs the identical operation, removing the objective standard, substituting the peer distribution, and publishing the result continuously.
Here the evidence refuses to line up behind the obvious conclusion. A 2023 systematic review screened 548 articles and analysed 40, finding points, badges and rankings widely used and reporting a positive influence of gamification on motivation [47]. Damien Fleur and colleagues ran two randomised controlled interventions in higher education, showing students their current and predicted performance beside peers with similar goal grades, chosen to produce slight upward comparison [48]. The dashboard raised extrinsic motivation and achievement, with no effect on metacognition.
That contradicts nothing above. It is the opening distinction arriving again. Slight upward comparison against peers chasing the same goal is what students pick for themselves and benefit from. A fixed quota is the other thing, where your gain requires someone's loss.
Comparison is not the mechanism that damages. Negative interdependence is. The warning about leaderboards is narrower than usually made, and what matters is whether the ranking is attached to a scarce prize.
What To Do Instead
The evidence supports fewer changes than the literature proposes, and they hang together rather than sitting in a list.
Everything starts with restoring the standard Glaser named in 1963 and Bloom argued for in 1968. Grade against a stated criterion rather than the cohort, so a student scoring 85 percent gets an A whether they are best in the room or worst. That change makes the rest usable, because each finding below depends on the grade carrying information about the material rather than the ranking.
Butler's result then becomes available, and it is close to free. If you write comments, do not put a normative score beside them, because the score cancels the comment. Where a mark must be recorded, separating it in time from the feedback lets the student read the sentences first.
Group work depends on the same condition, which is why schools implement it badly. Cooperation only pays when the group shares a goal and each member is individually accountable, and without both, Slavin's reviews put the median effect near +0.07. Recent classroom studies find the properly structured version still working, including a 2024 study in Ethiopian university classrooms [49] and work connecting cooperative structures to friendship formation rather than only to grades [50].
Underneath all of it sits the goal structure the room broadcasts. Carole Ames showed in 1992 that a classroom's evaluation and recognition practices, rather than the student's personality, push a room toward mastery or performance goals [51]. Lisa Bardach and colleagues confirmed that transmission across 68 studies covering 47,975 students, each goal most strongly related to its own contextual counterpart [52]. What the room rewards publicly is the lever, which is why grading policy matters more than any exhortation to learn for its own sake. Our piece on mastery goals and performance goals covers how those orientations work in an individual learner.
A caution belongs here, since this article has argued that measured constructs deserve scepticism. Achievement goal measurement has a known problem of its own, with different labels attached to similar constructs across studies, so read the transmission finding as a direction rather than a precise quantity.
What the evidence does not support is the idea that removing grades solves the problem. Fleischmann's Swedish result stands against it. Students compare regardless. What a criterion-referenced system changes is not whether comparison happens, but whether the institution makes comparison the only available measure and then attaches consequences to it.
What This Means If You Are the One Being Graded
You did not choose the policy, so here is what the research does and does not license you to conclude.
If your course is curved, your grade is partly a fact about your classmates. That is the arithmetic of the policy, not a rationalisation, and Murphy and Weinhardt's work suggests your rank carries forward in ways unrelated to what you learned.
If you feel less capable in a strong programme than in a weaker one, that is the best-replicated finding in this article, measured across more than a million students in every country tested. It is a fact about the reference group, not about you.
The comparison instinct itself is not the enemy. Dijkstra's review found students choose to compare upward and often improve for it, and Natacha Boissicat and colleagues showed in 2021 that priming a student toward approaching success rather than avoiding failure changes the effect of comparison on self-evaluation [53]. What you compare for matters more than whether you compare. If the ranking has started to feel like the point of the course, our article on learned helplessness covers what repeated uncontrollable outcomes do to effort, and the piece on impostor syndrome deals with the gap between a record and a self-assessment.
The most useful move is the one the research keeps pointing at without saying it. Find an objective standard and use it. Past papers with published mark schemes, a competency checklist, end-of-chapter problems answered without help. Anything that tells you what you know rather than where you sit. That is the standard Festinger said you would prefer, and a curved course has taken it away, so you supply it yourself. Our guides to the testing effect and desirable difficulties cover that kind of self-measurement, the psychology of high-stakes exams deals with performing under evaluative pressure, and stress and memory explains what sustained evaluative threat does to recall.
What Is Not in Dispute
Plenty here is unsettled, and this article has named both sides each time it came up. Whether competition helps performance at all, where Murayama and Elliot found no net effect and the Johnsons dispute the analysis. Whether removing grades removes the comparison, where the Swedish evidence says it does not. Whether pass and fail costs anything academically, with San Diego's data against Mayo's. What any of it does to equity, where the findings point several ways at once.
The structure is not in dispute. A fixed quota makes one student's success another student's cost. Deutsch described what that arrangement produces in 1949, and most of what has been measured since is a footnote to him.
So the question worth asking about your own institution is not whether its grading is fair. It is what its grading makes students want. A room where the goal is to learn the material and a room where the goal is to finish above the person next to you can hold the same students, the same teacher and the same syllabus, and produce different people.
Frequently Asked Questions
Is grading on a curve fair?
It depends which practice is meant. Adding points to a badly calibrated exam takes nothing from anyone. A fixed quota capping how many students may receive an A means two students with identical knowledge can get different grades depending on who else enrolled. The research treats that as uninformative more than unfair, because a relative score says nothing about what was learned.
Can grading on a curve lower my grade?
Under a true fixed distribution, yes. If only a set share of the class may receive each letter, improving your own score does not guarantee a better grade, because position rather than performance decides it. Under the additive version it cannot. Ask whether the whole class could in principle receive an A.
Does competition actually improve student learning?
The question is malformed. Murayama and Elliot's 2012 meta-analysis found no net relation between competition and performance, because competition pushes students toward performance-approach goals which help and performance-avoidance goals which hurt, and the two cancel. Cooperative against competitive structures is clearer: across 148 studies and 17,000 adolescents, cooperation led on achievement at d = 0.46 and peer relationships at d = 0.48.
What are the alternatives to grading on a curve?
Criterion-referenced grading against a stated standard is the direct replacement, restoring the objective measure a curve removes. Standards-based, specifications and mastery approaches are variations on it. Pass and fail has the most institutional evidence, with better wellbeing and no consistent loss in academic outcomes, though medical program directors treat it as a selection disadvantage. One caution: removing grades does not remove comparison, and demand for a ranking reappears wherever a scarce prize does.




