Introduction

Somebody walks into the room. Before they have said anything worth hearing, you have already decided several things about them. Whether they are competent. Whether they are warm. Whether you would trust them with something that matters.

You did not sit down and work any of that out. It arrived.

That much is not controversial. What happens next is where the interesting problem lives. Once one of those judgements has landed, it starts colouring the others. If the first impression was good, the second and third and fourth tend to come out good as well, and they keep coming out good even when they are about things the first impression could not possibly tell you.

That is the halo effect, and you will find it defined in almost identical words on a dozen pages. A positive impression in one area spreads to unrelated areas. Attractive people are assumed to be kinder. Confident speakers are assumed to be more competent. A brand you like makes a product you have never used seem better.

All of that is true, as far as it goes. The trouble is that it stops exactly where the subject gets interesting.

Because the halo effect did not enter psychology as an observation about people. It entered as a complaint about a form. The paper everyone cites is called [1] A constant error in psychological ratings, and the man who wrote it was not announcing that we judge books by covers. He was reporting that a rating scale used by the United States Army could not tell four things apart, and he thought that was a serious problem with the scale.

A hundred and six years later, nobody has fully fixed it.

This article follows that argument from 1920 to now. It is not a list of examples. It is one question, asked early and answered late: when a rater says a person is good at several things, how much of that agreement is distortion and how much of it is the person actually being good at several things? The field has fought about this in print for forty years, and the fight is not over. What follows is the story of that fight, what it settled, and what it did not.

Along the way it will tell you some things the popular version leaves out. That people deny this is happening while it is measurably happening to them. That training raters to avoid it has been shown to make their ratings less accurate. That in at least one large study the halo got weaker when faces were run through a beauty filter, which is the opposite of what almost everyone predicts.

The Paper Everybody Cites and Almost Nobody Reads

Edward Thorndike was an educational psychologist. In 1920 he published a short paper in the Journal of Applied Psychology reporting something odd in military ratings [1].

Commanding officers had been asked to rate their soldiers on separate qualities. Physical characteristics. Intelligence. Leadership. Personal qualities. Four different things, on four different scales, so four largely independent sets of numbers ought to come back.

They did not. The numbers agreed with each other far more closely than four independent qualities plausibly could. An officer who thought a man was physically impressive also thought he was intelligent, and a good leader, and a decent person. The four scales were behaving like one scale wearing four hats.

Read the title again. A constant error in psychological ratings.

Thorndike was not delighted. He was worried. A constant error, in the language of measurement, is a bias that does not cancel out when you average across many observations, which is what makes it far more dangerous than random noise. Random noise you can drown with more data. A constant error just gets more confident.

His concern was practical and it was about the instrument. If the Army was going to select and promote people on the strength of these ratings, and the ratings could not separate a soldier's intelligence from his bearing, then the ratings were not doing the job they had been built for.

That framing gets lost almost everywhere the effect is now described. The halo effect is presented as a discovery about human nature. It began life as a bug report.

Holding onto the original framing matters, because it tells you where to look for the argument. If you think the halo effect is a fact about people, the natural question is how big it is and how to stop it. If you think it is a problem with a measuring instrument, the natural question is much harder: how do you tell the instrument's error apart from the thing it is measuring?

That second question is the one this article is about.

What a Rating Sheet Is Supposed To Do

Picture the form. Four boxes, four scales, one soldier.

The whole design rests on an assumption so ordinary that nobody states it. The assumption is that the four qualities are separate, so a rater can consider them one at a time and give four independent answers.

Now consider what it means when the four answers agree.

One possibility is that the rater has formed a single global impression, positive or negative, and is filling in four boxes from it. That is distortion. The form asked four questions and got one answer copied four times.

The other possibility is more awkward. Maybe some soldiers really were better across the board. Health, education and confidence are not independent in real populations. A person who is fit, well fed, well schooled and comfortable in a hierarchy will genuinely tend to score higher on several of those scales at once, not because the rater is confused, but because those things travel together in the world.

If that is what is happening, then correlated ratings are correct. Removing the correlation would make the ratings worse.

You now have two explanations for exactly the same pattern in exactly the same numbers, and no way to tell them apart from the numbers alone. That is the trap. Everything else in this article is a hundred years of people trying to climb out of it.

The Traits That Reorganise Everything Else

The measurement problem sat mostly inside industrial psychology for two decades. What brought the halo into the social psychology mainstream was a different kind of experiment.

Solomon Asch gave people a short list of adjectives describing an imaginary person and asked what that person was like [2]. The lists were nearly identical. In one version the person was intelligent, skilful, industrious, warm, determined, practical and cautious. In another, the word warm was replaced with cold and everything else stayed the same.

One word. Everything downstream moved.

The warm version was described as generous, wise, happy and good natured. The cold version, built from the same six remaining adjectives, was not. Asch's conclusion was that impressions are not built by adding traits together like a shopping list. Some traits are structural. They tell you what kind of person you are looking at, and every other trait then gets read through them.

Harold Kelley took the same idea out of the laboratory and put it in front of a real class [3]. Students were given a short biographical note about a guest lecturer before he arrived. Half the notes described him as rather warm. Half described him as rather cold. The lecturer then taught the same session to everyone.

The students who had read warm liked him more, rated him better, and were more likely to join the discussion. Same man. Same twenty minutes. A different adjective read in advance.

That result has been replicated in evaluation settings since, including work looking directly at how warm and cold framing shifts ratings of teaching effectiveness [4]. It is one of the reasons student evaluations are treated with such care by people who study measurement, and with so little care by everybody else.

Notice what has happened to the problem. Thorndike found ratings that would not separate. Asch and Kelley showed how they get stuck together in the first place. A single early piece of information sets a frame, and the frame does the rest of the work quietly.

Warm amber glow illuminating pale stone slabs in dark composition.

What Is Beautiful Is Good

In 1972 Karen Dion, Ellen Berscheid and Elaine Walster published a paper with a title so quotable it became the name of a whole research area [5]. What is beautiful is good.

The design was simple. People looked at photographs of strangers and then answered questions about them. Not questions about how they looked. Questions about their personalities, their likely marriages, their likely careers, their likely happiness.

The attractive strangers came back better on nearly everything.

Two years later Landy and Sigall ran the version that makes the point harder to argue with [6]. They gave people an essay to evaluate. The essay was attached to a photograph of its supposed author. Same essay, different photograph, different scores. The title they chose says it plainly. Beauty is talent.

This is where a lot of coverage goes wrong, so it is worth being careful. What these studies show is that attractive people are rated more favourably on traits that have nothing to do with appearance. They do not show that attractive people possess those traits. The rating is the finding. The person is not.

That distinction sounds pedantic. It turns out to be the hinge the entire modern literature swings on, and we will come back to it with numbers.

The Part That Should Bother You

Here is the study that changed how I read all of this.

In 1977 Richard Nisbett and Timothy Wilson had students watch a videotaped interview with an instructor who spoke with a Belgian accent [7]. Some saw him behave warmly and pleasantly. Others saw the same instructor behave coldly and rigidly. Then everyone rated him.

Not just on likeability. On his appearance, his mannerisms, and his accent.

Those three things were physically identical across conditions. The accent did not change. Yet the students who had seen the warm version rated his appearance, his mannerisms and even his accent as more appealing. The students who had seen the cold version rated the same features as irritating.

So far this is a clean halo demonstration. The unsettling part is what happened when the researchers asked the students about it.

They denied it. Participants did not report that liking the man had made his accent sound better. Some of them reported the arrow running the other way, saying that their reaction to his mannerisms had shaped how much they liked him. That is not a small misremembering. It is a confident account of your own reasoning that runs backwards to what the data show.

Nisbett and Wilson wrote a companion paper the same year making the general case [8]. When you introspect on why you reached a judgement, you are not reading off a record of the process. There is no such record available to you. You are constructing a plausible story after the fact, from the same folk theories everybody else uses.

Sit with what that means practically. The most common piece of debiasing advice in the world is some version of check yourself. Notice when you are doing it. Be aware of your biases.

The people in this study were doing exactly that, and they got it wrong. If you want the longer version of why self-inspection is such a poor instrument, it is the same failure described in the illusion of knowing, where confidence in a judgement turns out to be almost independent of whether the judgement is any good.

Yes

No

One salient cue

Fast global impression

Unrelated traits rated

Asked to explain?

Plausible story invented

Judgement acted on

Influence sincerely denied

The diagram makes it look tidier than it is. In practice the whole left side happens in well under a second, which is the subject of a later section, and the denial is not a lie. It is what honest introspection returns.

Ubiquitous

By 1981 there was enough evidence for William Cooper to write a review with a one word title [9]. Ubiquitous halo.

His case was that halo is not an occasional problem in badly designed rating forms. It is everywhere, it inflates the correlations between rated dimensions almost universally, and researchers who use ratings as data are routinely reporting relationships that are partly an artefact of how the ratings were collected.

If Cooper is right, the damage is not limited to bad hiring decisions. It reaches into any field that measures anything by asking a person to rate it. Which is most of the social sciences, all of performance management, most of medical education, and a great deal of consumer research.

That is a serious claim. It also set up the counterattack.

The Distinction the Internet Never Mentions

Go and read any of the pages currently ranking for this topic. You will find the definition, a set of examples, the horn effect as its evil twin, and some advice.

You will not find this.

There are two different things buried inside a correlated set of ratings, and the literature has names for them. True halo is the part of the correlation that reflects reality, because the rated qualities genuinely go together in the person being rated. Illusory halo is the part contributed by the rater, because a global impression bled into every box.

Only the second one is an error.

This is not a philosophical nicety invented to complicate things. It is the central practical problem, and it has been argued about continuously. Balzer and Sulsky examined the state of halo research in performance appraisal and concluded that much of it rested on assumptions nobody had tested [10]. Fisicaro looked at the supposed relationship between halo and accuracy and found it was not the simple inverse everyone assumed [11]. More halo did not reliably mean less accuracy.

TermWhat it meansIs it an errorEveryday example
Halo effectOne positive impression spreads to unrelated judgementsDepends entirely on the next two rowsA confident speaker is assumed to be well prepared
Horn or devil effectOne negative impression spreads the same waySame answer as aboveA single spelling mistake makes a whole proposal seem sloppy
True haloThe part of the agreement that is real because the qualities really do go togetherNoA well organised colleague genuinely is also punctual
Illusory haloThe part the rater added because they had already formed a global viewYesScoring somebody high on creativity because they are likeable

Read the bottom two rows again, because that is the whole problem. When you look at a set of ratings you cannot see which row you are in. Both produce the same shape in the data. The correlation does not carry a label.

The Result That Should Have Ended the Story

In 1990 Barry Nathan and Nancy Tippins published a field study that is very rarely mentioned outside the specialist literature, and it deserves to be [12].

They looked at how the amount of halo in performance ratings affected the results of test validation. If halo is an error, adding it to ratings should degrade the relationship between a selection test and the performance it is meant to predict. Error is noise. Noise weakens relationships.

That is not what they found. Ratings carrying more halo produced better validation results, not worse.

Think about what a genuine measurement error is supposed to do. It should push your numbers around at random and blur real relationships. If adding more of something makes your predictions sharper, the something you are adding is carrying information.

The most natural reading is that a lot of what gets counted as halo is true halo. When a supervisor forms a global impression, that impression is not made of nothing. It is often a decent summary of how good somebody is at their job overall, which is exactly what a selection test is trying to predict.

One field study does not settle a literature. But it does something important. It makes the confident version of the story, in which halo is simply an error to be removed, impossible to keep asserting without evidence.

Are We Even Measuring It

Three years later Kevin Murphy, Robert Jako and Rebecca Anhalt wrote the paper that took the whole edifice apart [13].

Their argument went after the assumptions rather than the findings. They challenged the ideas that halo is common, that it is always an error, that the standard statistical measures of halo actually measure halo, and that removing it improves the quality of ratings. On their reading, decades of research had been carefully quantifying something without establishing what it was.

Murphy and Anhalt had already published a related result that is worth its own sentence [14]. They asked whether halo is a property of the rater, of the people being rated, or of the specific behaviours observed. If halo were a stable characteristic of certain raters, you could screen for it, train it out, or weight around it. It did not behave that way. It shifted with what was being rated and by whom.

That finding demolishes a very popular idea. There is no such thing as a reliably halo prone person you can identify and correct.

Solomonson and Lance later examined the relationship between true halo and halo error directly and found the two are entangled in ways that make clean separation extremely hard in practice [15]. Which is the same wall Thorndike hit, seventy seven years later, with far better statistics.

The Closest Thing to an Answer

The best available resolution is a meta-analysis that most people writing about the halo effect have never cited.

Chockalingam Viswesvaran, Frank Schmidt and Deniz Ones built a database integrating ninety years of studies reporting how rated job performance dimensions correlate with each other [16]. Then they asked the question directly. After you strip out halo and three other sources of measurement error, is there anything left?

There was. A general factor in rated job performance survived and accounted for sixty percent of total variance. Something more than one rater's impression is running through those numbers.

Keep the distinction from earlier in place, though. That is still a result about ratings, not a direct measurement of people. What it establishes is narrower and still important: the agreement between dimensions is not purely invented.

But the correction they had to apply was not small. Halo inflated the correlations between dimensions rated by the same person by thirty three percent for supervisors and sixty three percent for peers.

Both halves are true. There is a real general factor, and there is substantial rater inflation sitting on top of it, and peer ratings carry roughly twice the inflation supervisor ratings do.

Three separate findings from one 90-year meta-analytic databaseGeneral factor varianceSupervisor inflationPeer inflation7065605550454035302520151050Percent

Those three bars are three separate findings and they do not add up to anything. The first is how much of job performance is genuinely general. The second and third are how much same rater correlations were inflated by halo before correction. Putting them side by side is only meant to show you the scale of each.

That is roughly where the argument stands. Not resolved. Quantified.

The Problem Did Not Stay in Psychology

Once you know the shape of this problem you start seeing it in fields that do not use the word halo at all.

William Hoyt wrote a methods paper on rater bias in psychological research generally, working through when it distorts conclusions and what can be done about it [17]. Thomas Feeley made a similar case for communication research [18]. Gary Pike wrote about it in higher education outcomes research, and reached for Thorndike's phrase in his own title [19]. The constant error of the halo.

Statisticians are still building tools for it. Jin and Chiu built a mixture Rasch model specifically to detect illusory halo in rating data, which is not something you do about a solved problem [20].

Any time a number in a study came from a person filling in a scale, some of this is baked into it. That includes a great many numbers you have read confidently reported.

Clear water glass with indigo color swirling through it.

Meanwhile the Attractiveness Literature Kept Winning

While industrial psychologists were arguing about whether halo was even measurable, the attractiveness branch of the field was accumulating one of the sturdier bodies of evidence in social psychology.

Alice Eagly and colleagues published a meta-analytic review in 1991 with a title that contains its own qualification [21]. What is beautiful is good, but. The stereotype was real, and it was strongly moderated. The effect was largest for judgements of social competence. It was much weaker for judgements of integrity and concern for other people.

That is a more interesting finding than the slogan it corrects. People are not simply assuming attractive strangers are better in every way. The inference has a shape. Beauty buys you a big advantage on how socially capable you seem and a much smaller one on whether you seem honest.

Nine years later Judith Langlois and colleagues published a paper containing eleven separate meta-analyses [22]. It is the widest treatment of the topic anyone has produced, and four of its conclusions matter here.

Raters agree about who is attractive, both within cultures and across them. Attractive children and adults are judged more positively. Attractive children and adults are treated more positively, including by people who already know them well. And attractive children and adults do display somewhat more positive behaviours and traits.

That fourth one is the true halo problem walking straight into the attractiveness literature. If attractive people really do behave somewhat differently, perhaps because they have been treated differently their whole lives, then favourable ratings of them are not purely an error either. The bias may be partly self fulfilling rather than purely imaginary, which is a much less comfortable finding than the usual framing.

One Hundred Milliseconds

The speed of this thing is the part most people underestimate.

Janine Willis and Alexander Todorov ran five experiments manipulating how long people saw an unfamiliar face before judging it [23]. Judgements made after a hundred milliseconds correlated highly with judgements made with unlimited time, across attractiveness, likeability, trustworthiness, competence and aggressiveness.

Giving people more time did not change the verdict. It raised their confidence in it.

A hundred milliseconds is less than a blink. There is no deliberation in that window and nothing you would recognise as thinking. The judgement is already formed by the time you have any sense of having looked.

Todorov's group had already shown these snap judgements reaching into consequences that matter [24]. Competence inferred from photographs alone, by people who did not recognise the candidates, predicted the outcome of United States congressional races better than chance. In the 2004 Senate races the figure was 68.8 percent, and the judgements were also linearly related to the eventual margin of victory. Later work took the same finding apart specifically in terms of attractiveness and candidate evaluation [25].

Elections are supposed to be the deliberative end of human decision making. Voters read positions. They watch debates. And a still photograph shown for a second to a stranger in another state carries real predictive signal about the result.

Beauty Did Not Add Noise. It Hid a Signal.

Now the study that changes what the bias actually costs you.

Sean Talamas, Kenneth Mavor and David Perrett showed people the faces of a hundred university students and asked them to judge academic performance [26]. They also knew the students' real academic performance, so they could check the judgements against something.

Two findings, in this order.

First, there were genuine cues. People had some real ability to read academic performance from a face, above chance. That is not mystical. Sleep, health, stress and grooming leave visible traces, and earlier work had already found some valid signal linking facial appearance to measured intelligence beyond the attractiveness halo [27].

Second, the attractiveness halo overwhelmed it. Raters leaned so heavily on how attractive a face was that they lost access to the weaker, more accurate cues they were otherwise capable of using.

That is a different and worse story than the usual one. The standard account says bias adds an unfair advantage. This says bias destroys information you already had. The raters were not simply being unjust to unattractive students. They were being made worse at a task they could partly do.

That was a photograph and a judgement about grades, so do not stretch it further than it goes. What it does show is a trade that cost the raters real accuracy. Attractiveness is loud and instantly readable. The cue carrying the real information was quiet. Loud won.

The Classroom

The grading evidence is the best quantified consequence in this whole field, so it deserves a real number.

John Malouff and Einar Thorsteinsson pooled experimental studies of grading bias in a meta-analysis covering twenty three analyses from twenty studies and a total of one thousand nine hundred and thirty five graders [28]. In each study graders were given some piece of information about a student that had nothing to do with the work in front of them. The overall between groups effect size was g equals 0.36.

That is not enormous and it is not trivial. It is roughly the difference between a solid pass and a slightly better one, applied systematically, to the same piece of work, on the basis of who the grader thought wrote it. Moderator analyses found no significant difference between school work and university work.

The cleanest real world test is not an experiment at all. Rey Hernandez-Julian and Christina Peters used a natural comparison, looking at how student appearance related to grades in courses where the instructor could see the students against courses delivered online where they could not [29]. Removing the cue removes much of the problem, which is a design lesson rather than a personal one.

The evaluation runs in the other direction too. Nalini Ambady and Robert Rosenthal showed people silent video clips of teachers lasting seconds, with no sound and no content, and used those ratings to predict the teachers' end of term evaluations from students who had sat through an entire course [30]. The thin slices predicted the full term ratings.

Whatever end of term teaching evaluations measure, a substantial part of it was decided in the first thirty seconds. Classroom observation ratings by school principals show the same family of rater effects when analysed properly [31].

Pale river stones on dark slate, one glowing warmly.

The Interview Room

Hiring is where the evidence about fixing this is strongest. That is convenient, because hiring is also where a bad judgement is hardest to walk back.

The unstructured interview is the standard method almost everywhere. A conversation. Some questions. A judgement at the end. It is also, by a wide margin, the format most exposed to everything in this article, because an impression formed in the first minute has forty more minutes to justify itself.

Julia Levashina and colleagues reviewed the structured employment interview literature in detail [32]. Structure means the same questions in the same order, scored against defined criteria, ideally by more than one person. It consistently outperforms the conversational version.

So why does anybody still run unstructured interviews?

Jason Dana, Robyn Dawes and Nathanial Peterson answered that with an experiment whose title says it all [33]. Belief in the unstructured interview: the persistence of an illusion. Interviewers keep believing in their own judgement even when shown evidence that it is not working. That belief is itself a halo of sorts, sitting on your own impression of your own skill.

There is something almost funny in that result and something bleak. The skill people defend hardest is the one with the least evidence behind it.

It shows up in professions that pride themselves on scepticism. Marc Eulerich and colleagues looked at auditors making fraud risk judgements and found attractiveness influenced them [34]. Auditing is a field built entirely around not taking things at face value.

Recent experimental work with government employees found the same appraisal bias in a public sector setting, where evaluations feed directly into promotion and pay [35]. None of these people were careless. Halo does not require carelessness. That is what makes it durable.

The Courtroom and the Swindler

The legal studies are the ones most often mangled in summary, so here is the careful version.

Michael Efran ran a simulated jury task in 1974 in which students judged a defendant's guilt and recommended punishment [36]. Attractiveness affected both. Note the word simulated. These are students in a room with a case file, not jurors in a courtroom, and the whole literature carries that limitation.

The result worth knowing is the one that reverses. Harold Sigall and Nancy Ostrove varied both the attractiveness of an offender and the nature of the crime [37]. For a burglary, attractiveness helped, in the expected direction. For a swindle, it hurt. Recommended sentences went up.

The reading is straightforward once you see it. When the crime was one where good looks were the instrument of the offence, attractiveness stopped being a mitigating quality and became evidence of how the crime was committed.

That single reversal is worth more than fifty confirmations, because it tells you the halo is not a blanket bonus applied to attractive people. It is an inference, and inferences are sensitive to context. More recent work has examined how the relationship between facial attractiveness and perceived guilt varies across crime types, which is the same question with better methods [38].

Anyone who tells you attractive defendants always get lighter sentences has not read the second study.

The Horn Effect Costs More

The negative version has two names and neither is flattering. The horn effect. The devil effect.

Mechanically it is the same process running with the sign flipped. One unfavourable impression spreads outward and starts colouring unrelated judgements. Mary Radeke and Anthony Stahelski demonstrated both directions with facial photographs, showing that age and gender stereotypes could be pushed either way [39]. Recent work has examined the devil effect in the context of sexual crime allegations, where the stakes are as high as they get [40].

The reason to treat this half seriously is who it lands on.

Mariola Paruzel-Czachura and colleagues studied first impressions of faces with scars and facial palsies, measuring warmth, competence and humanization [41]. That last word is the one to stop at. The question was not only whether people rated these faces lower on friendly traits. It was whether they attributed less full humanity to them.

Nobody in that situation chose their face. The halo effect discussed as a curiosity about attractive people is the same mechanism that quietly penalises people with visible facial difference in every room they walk into.

There are also conditions where attractiveness costs you. Work on trust decisions found that beauty is not always a perk, and that attractiveness interacts with perceived motives rather than simply raising them [42]. If you look like you could have anything you want, being trusted with something is not automatic.

Bare stone chamber illuminated by a deep red glass pane.

Where Does It Actually Come From

Four explanations compete, and the honest answer is that they are not fully separable yet.

The first is structural. Seymour Rosenberg, Carnot Nelson and P.S. Vivekananthan mapped how impressions of personality are organised and found they collapse onto a small number of dimensions, essentially a social good to bad axis and an intellectual good to bad axis [43]. If your mental filing system for people has two drawers, then learning one thing about somebody puts them in a drawer, and everything else in that drawer comes along. David Schneider's review of implicit personality theory laid out how much of impression formation runs on these assumed structures rather than on evidence [44].

The second is affective. You feel something about the person, and the feeling colours the individual judgements. Joseph Forgas tested this by manipulating mood and measuring the size of the halo, and it moved [45]. His title captures the effect neatly. She just doesn't look like a philosopher. If mood changes the size of the bias, then part of the bias is made of mood, which connects it directly to how emotion shapes what you remember and how you weigh it.

The third is associative learning. Marine Rougier and Jan De Houwer have argued that halo behaves in some ways like evaluative conditioning, and importantly that it updates when you get new information rather than sitting there permanently [46]. Their work with colleagues explored the links between impression formation and conditioning in more depth [47]. This matters because it is the most optimistic finding in the article. A halo that updates is a halo that new evidence can shift.

The fourth is inferential. You are not colouring, you are reasoning, from a belief that traits genuinely correlate in the world. Research finding that competence spills into perceived morality fits this [48], as does work on the moderators of the liking bias in moral character judgements [49].

The inferential account is stronger than it first looks, because sometimes traits genuinely do not travel together. Studies of compensation effects find warmth and competence trading off against each other rather than haloing [50]. Work on the limits of the primacy of morality shows the same boundaries [51]. A review of the stereotype content model across psychology and marketing maps where these dimensions do and do not cohere [52]. If people were simply smearing positivity everywhere, that trade off should not exist.

Trait observed

Structural good-bad axis

Affective colouring

Associative pairing

Inferred trait correlation

Correlated ratings

Four routes, one output. That is exactly why the mechanism question is hard: the data at the bottom of the diagram look the same whichever path produced it. Clare Sutherland and Andrew Young's integrative review of trait impressions from faces is the best current attempt to hold the perceptual and conceptual accounts together in one framework [53].

The Answer Hiding in the Dictionary

In 2024 Chris Westbury and Daniel King published a paper whose title is a deliberate echo [54]. A Constant Error Revisited.

Their proposal is that a chunk of the halo effect is not in the perceiver at all. It is in the language.

Trait words are not independent of each other. Words that appear in similar contexts carry similar connotations, and English has spent centuries braiding its vocabulary for people into clusters. If you are asked whether a generous person is also likely to be honest, you may not be reasoning about people at all. You may be reasoning about words.

They tested it. In the first study thirty nine participants judged how likely it was that pairs of traits would occur together, across a hundred and twenty six trait pairs. The similarity between the two words in a semantic vector space predicted those judgements at a cross validated r squared of 0.19, and adding two further word similarity measures pushed the variance explained to forty five percent. A second experiment with forty different participants confirmed the word pairs were not simply synonyms, with an average judged similarity of 40.8 out of 100.

Forty five percent of the variance in whether people think two traits go together is predictable from how the words behave in text.

That does not dissolve the halo effect. Language reflects real regularities as well as inventing them. But it does raise a possibility that had not been tested properly. Part of what Thorndike measured in 1920 may have been living in the rating form's vocabulary rather than in the officers' heads. This is one laboratory result about trait words, with 39 people and then 40 more, and it is not a verdict on a century of field research. It is a good enough idea to want tested outside a laboratory.

What Happens When You Tell People To Stop

Here is the practical section, and it starts with two results that most advice ignores.

Walter Borman tested instructions to avoid halo error, telling raters explicitly about the problem and asking them not to do it [55]. The instructions changed the ratings. They did not reliably improve them.

Then John Bernardin and Earl Pence tested full rater error training, the formal version organisations still buy today [56]. Their title is the finding. Effects of rater training: creating new response sets and decreasing accuracy.

Read that once more. The training decreased accuracy.

What happened is worth understanding, because it generalises. Raters were taught that giving similar scores across dimensions is an error. So they stopped doing it. They started spreading their scores out, because spread scores looked correct.

But some of that agreement was true halo. Some people really are good at several things. By training raters to avoid the pattern, the researchers trained them to avoid a pattern that was partly accurate, and the ratings drifted further from reality while looking more sophisticated.

This is the single most useful thing in the article and it applies well beyond rating forms. If you correct for a bias without knowing how much of the pattern is real, you can overshoot. You get a different error and more confidence in it.

ApproachDoes it workWhat the evidence found
Tell raters about the biasNot reliablyInstructions to avoid halo changed ratings without dependably improving them
Formal rater error trainingIt can backfireTraining created new response sets and decreased accuracy
Structured interview with fixed questions and scored criteriaYesConsistently outperforms unstructured conversation across the review literature
Rate specific behaviours rather than global qualitiesYesReducing the influence of prior performance expectations improved behavioural ratings
Remove the cue entirelyYesComparing in person with online courses shows what disappears when appearance is not visible

What Actually Works

The pattern in that table is not subtle. Nothing that relies on the rater trying harder works. Everything that changes the structure of the task does.

Boris Baltes and Christopher Parker showed that reducing the influence of prior performance expectations on behavioural ratings is achievable, but through changing what raters are asked to do rather than through exhortation [57]. Structured interviews work for the same reason [32]. Blind marking works for the same reason. Comparing courses where the instructor can and cannot see the student works for the same reason [29].

Ask about specific behaviours instead of overall quality. Fix the questions in advance. Score each dimension before you see the next one. Have more than one person rate. Where you can, remove the irrelevant cue from the process entirely.

None of that requires anybody to become less biased. It requires the process to give the bias less to work with. If you want the general version of that idea, it is the same reason structured self monitoring beats good intentions in learning: judgement improves when it is scaffolded, not when it is scolded.

There is one genuinely hopeful finding to add. The halo updates. Rougier and De Houwer found the attractiveness halo changes when new and relevant information arrives [46]. First impressions are sticky. They are not sealed.

Did It Survive the Replication Crisis

You should ask this about any famous finding in social psychology now, and if you arrived here after reading about the Dunning-Kruger effect you have already seen how differently these stories can end.

The short answer is that the halo effect held up. It is among the better replicated phenomena in the field, with a century of confirmation behind it and cross-cultural evidence added in the last few years. Carlota Batres and Victor Shiramizu examined the attractiveness halo across cultures [58].

A preregistered study by Tobias Kordsmeyer and colleagues went further and took the effect off the face entirely [59]. They photographed and 3D body scanned 165 German men, then had 123 German and 100 Japanese observers rate attractiveness, prosociality, health and physical dominance from faces and bodies separately. Strong attractiveness halo effects appeared for both faces and bodies, and the results were largely consistent across the two observer groups. Other work has asked the generalisability question directly, noting how much of the literature rests on Western adult samples [60].

Now the part a careful article has to include.

Recent studies have gone looking for halo effects in specific settings and failed to find them. Alexis Makin, Autumn Taylor and Lauren Macpherson gave 302 undergraduates a realistic academic malpractice vignette alongside a photograph of an attractive or unattractive student and asked about guilt, appropriate punishment and seriousness [61]. There was no evidence for halo effects.

Robin Kramer and colleagues tested whether the order in which you see somebody's photographs anchors your impression of their attractiveness, and in an experiment with 301 participants found the effect was minimal [62].

The most instructive null is a prediction that failed. One account of why attractive candidates win votes says attractiveness is read as a cue to health, so a disease threat should push voters towards attractive leaders. A pandemic is about as clean a test of that as the world will ever hand anybody. Lasse Laustsen and Asmus Leth Olsen ran six of them, using two nationally representative Danish surveys of 3297 people, one at the outbreak and one a year later [63]. Disease threat did not raise the preference for attractive leaders. The prediction did not survive.

None of that overturns the meta-analyses. What it does is mark the boundary. The halo effect holds up in the aggregate and is unreliable in any particular instance, which is exactly what a moderate sized effect with lots of moderators looks like from close up. Anybody who tells you it will definitely happen in a specific situation is overselling it, and so is anybody who tells you it is a myth.

The Filters Made It Weaker

I expected the modern evidence to show the halo getting stronger. Curated images, filtered faces, a decade of visual self presentation. The prediction almost writes itself.

Aditya Gulati and colleagues ran the study that tests it, and the result went the other way [64]. They recruited 2748 participants to rate facial images of 462 distinct individuals, each face appearing in an original version and in a version processed with an AI beauty filter.

The headline finding was the expected one. The same people were rated significantly higher on attractiveness and on unrelated traits, including intelligence and trustworthiness, in the beautified condition.

The second finding was not. The halo effect itself was weaker in the beautified condition. The authors argue this helps resolve conflicting results in the earlier literature, and they raise the possibility that filters could dampen the bias rather than amplify it.

A plausible reading is that when everybody looks polished, polish stops carrying information. A cue that is universally present is a cue that discriminates nothing. That is a mechanism worth watching as image processing keeps spreading.

Two newer questions follow straight from that. Robin Kramer compared ChatGPT's judgements of social traits from face photographs against human judgements [65], which asks whether systems trained on human data inherit human halos. And during the pandemic, work found that face masks reduced perceived trustworthiness [66], a reminder that hiding part of a face does not switch the machinery off. It just changes what it feeds on.

It Was Never Only About Faces

The face is where the research concentrated, and that is a fact about method rather than about the mind. Faces are easy to standardise in a laboratory. Photographs can be matched for lighting and cropping and shown in a fixed order, and nothing else about a person is that convenient.

That convenience has shaped what people think the halo effect is. Take the face away and it does not stop.

A voice carries none of the information a photograph does, and perceived vocal traits alone still shift how much money people will risk on a stranger in an economic game [67]. Whatever is being read there, it is not specifically visual. A trait you already believe about somebody will colour how you read something as small as their eye contact [68], and a coworker who simply seems happy picks up a halo across judgements of their other qualities [69]. It does not have to be a face or even a feature. It only has to be salient and arrive first.

Leslie Zebrowitz has spent a career on the part that is easiest to skip, which is whether the inference is ever right. With Robert Franklin she examined how the attractiveness halo and the babyface stereotype behave in older and younger adults [70]. Earlier work with colleagues went after perceived honesty and real honesty across the lifespan [71]. That second design is rarer than it should be. Most studies measure the inference and never check it against anything.

Her question is the one worth carrying out of this section. Not whether people infer, but whether the inference is any good. Part of what makes that so hard to answer is that the judgement will not hold still. How funny you find somebody shifts with your relationship to them and with the format you hear them in [72]. A verdict that moves with the setting is not simply right or wrong about the person, which is why accuracy is a harder thing to measure here than bias.

There is a connection here to how faces are handled generally. The reason a face dominates your first impression is partly that face processing is fast, automatic and enormously well resourced, which is also why you remember a face and lose the name attached to it. The fast system delivers a verdict before the slow one has read the question.

The Halo Around Things

The same mechanism works on objects, and the commercial world found this out early.

Pierre Chandon and Brian Wansink documented health halos from restaurant health claims, where a claim on the front of a menu led people to underestimate calories and to be more willing to add side dishes [73]. Two independent groups have since found the same pattern with different products. Health halo effects have been shown from product titles and nutrient content claims [74], and recent work has found an organic halo in which the same food is judged lower in calories when labelled organic [75].

Food is a useful test case, because taste is supposed to happen in your mouth.

The one that should worry professionals is a sensory science study. Trained descriptive analysis panels, people who have been taught to evaluate specific attributes independently, showed colour halo and horn effects when rating carbonated beverages [76]. If expert panels who have been explicitly trained to isolate attributes still let colour bleed into flavour ratings, an untrained person filling in a five point scale has no chance at all.

The same thing happens where there is no physical product at all. Attractive avatars in a hospitality setting shifted both what people chose to eat and how satisfied they were with it [77]. Attractive reviewers are more persuasive than unattractive ones writing exactly the same review [78].

Sit with those two for a second. Neither involves a person you could ever meet. One is a drawing. The other is a thumbnail beside a block of text you could have read with your eyes shut. The machinery does not need anybody in the room with you.

The commercial version keeps going from there. The attractiveness of a clothing model changes how the clothing itself is evaluated [79], and on online platforms the facial display chosen for a listing has been linked to how many viewers went on to convert offline [80].

Every one of those is Thorndike's problem wearing commercial clothes. One salient attribute is doing work that belongs to a different attribute, and the person doing the rating cannot feel the substitution happening.

Pale ceramic vessels on a dark shelf, softly illuminated.

A Hundred Years of Arguing About One Correlation

Laid out in order, the shape of the argument is easier to see than in any summary of it.

1920
Thorndike reports a constant error in Army ratings
1946
Asch shows one word reorganises a whole impression
1950
Kelley takes warm and cold into a real classroom
1972
Dion Berscheid and Walster publish what is beautiful is good
1977
Nisbett and Wilson find the influence is sincerely denied
1981
Cooper argues halo is ubiquitous in rating research
1990
Nathan and Tippins find more halo gave better validity
1993
Murphy and colleagues ask if halo is measurable
2000
Langlois and colleagues publish eleven meta-analyses
2005
Viswesvaran and colleagues separate the real general factor
2016
Talamas and colleagues show beauty hides a real cue
2024
Gulati and colleagues find the halo weaker under beauty filters
2024
Westbury and King locate part of the effect in language

The middle of that list is the part nobody outside the specialist literature ever sees. Between 1981 and 2005 the field ran a serious, quantitative, bad tempered argument about whether the thing had been correctly defined, and the resolution was not a winner. It was a decomposition.

What Is Settled and What Is Not

Settled, and you can say these plainly.

Ratings of separate attributes covary far more than the attributes themselves do. Attractive people are rated more favourably on traits unrelated to appearance, confirmed across two large meta-analytic reviews. Impressions form within about a hundred milliseconds and extra time raises confidence more than it changes the verdict. People do not have introspective access to the influence and will deny it in good faith. Grading is measurably affected, at g equals 0.36 across 1935 graders. The horn effect is the same mechanism with the sign flipped.

Not settled, and anybody claiming otherwise is going beyond the evidence.

How much of the correlation is real. Cooper says the distortion is pervasive, Murphy and colleagues say we may not be measuring it properly, Nathan and Tippins found halo carrying useful signal, and Viswesvaran and colleagues put numbers on both halves without dissolving the disagreement.

Which mechanism dominates. Structural, affective, associative and inferential accounts all have support, they are hard to separate empirically, and the newest proposal puts a large slice of it in the vocabulary rather than the perceiver.

Whether it is growing or shrinking in a filtered visual culture. The one large study built to test that found it weaker under beautification, which is one study going against a widespread assumption.

Whether awareness helps at all. The evidence that structure helps is good. The evidence that self monitoring helps is poor, and the evidence that formal error training can hurt is direct.

Conclusion

The halo effect is usually presented as a warning about first impressions, with an implied remedy: know about it, watch for it, do better.

The evidence does not really support that remedy. Nisbett and Wilson's participants were introspecting sincerely and got the direction of their own reasoning wrong. Bernardin and Pence's trained raters became less accurate, not more. Watching for it is not nothing, but it is much weaker than it sounds, and it is not what fixed the problem anywhere it has actually been fixed.

What worked was structure. Fixed questions. Defined criteria. Scoring one dimension at a time. More than one rater. Taking the irrelevant cue out of the room where that was possible.

There is a second and stranger conclusion, and it is the one Thorndike would recognise. Part of what looks like bias is not bias. Real people are not random collections of independent traits, so some of the agreement in your judgements is your perception working correctly. The uncomfortable part is that you cannot tell, from the inside, which part you are in. Neither can the field, after a hundred and six years of trying, and the same difficulty shows up whenever a belief and the evidence for it get tangled, which is the territory confirmation bias covers.

That is a less satisfying ending than five steps to a fairer mind. It is also more useful. The next time you are certain about somebody you met ten minutes ago, the right question is not whether you are biased. You cannot answer that one. The right question is what your impression is actually built from, and whether the process that produced it gave you anything to work with beyond the loudest cue in the room.

An impression you can trace is worth something. One that simply arrived, complete and confident, is exactly the one Thorndike was warning about. Your judgement of another person is reconstructed as much as it is recorded, in the same way memory rebuilds rather than replays, and a reconstruction always reflects the materials it was built from.

Frequently Asked Questions

What is the halo effect in psychology?

The halo effect is the tendency for one impression of a person to influence unrelated judgements about them. If your first impression is favourable, you are likely to rate that person more favourably on qualities the first impression could not tell you anything about, such as intelligence, honesty or competence. It was first reported by Edward Thorndike in 1920 in a paper called A constant error in psychological ratings, describing officers whose ratings of soldiers on separate qualities agreed with each other far more than four independent qualities plausibly could. The horn effect, sometimes called the devil effect, is the same process with a negative impression spreading instead.

Who discovered the halo effect and what did they actually find?

Edward Thorndike reported it in 1920, though the phenomenon had been noticed earlier. What matters is what he was arguing. His paper is titled A constant error in psychological ratings, and his concern was that a rating scale used to assess soldiers could not separate the qualities it was designed to measure. Officers who rated a soldier highly on physical qualities also rated him highly on intelligence, leadership and personal qualities. Thorndike treated this as a fault in the measuring instrument rather than as an interesting discovery about social perception. That framing has largely disappeared from popular accounts, and it is the framing that leads to the hardest question in the field: how do you tell rater distortion apart from qualities that genuinely go together?

Is the halo effect conscious or unconscious?

The evidence points strongly to unconscious. In a 1977 study by Richard Nisbett and Timothy Wilson, students watched the same instructor behave either warmly or coldly, then rated his appearance, mannerisms and accent. Those features were identical in both conditions, but they were rated differently depending on the manner they had seen. When asked, participants denied that liking the instructor had affected those ratings, and some reported the causal arrow running the opposite way. A companion paper the same year argued that people generally lack introspective access to the processes behind their judgements and construct plausible explanations after the fact. This is why advice to simply notice your own bias performs so poorly.

Has the halo effect replicated or did it fail the replication crisis?

It held up. Unlike several famous social psychology findings, the halo effect has a century of confirmation behind it, including two large meta-analytic reviews of the attractiveness stereotype and recent cross-cultural work covering faces and bodies with observers from different countries. But it is a moderate sized effect with many moderators, which means it is reliable in the aggregate and unpredictable in any single situation. Recent studies have looked for halo effects in specific settings and failed to find them, including one with 302 participants judging an academic malpractice case beside an attractive or unattractive photograph. Both things are true. The effect is real, and it does not appear on demand.

How do you reduce the halo effect in hiring or grading?

Not by trying harder. Studies that instructed raters to avoid halo error changed their ratings without reliably improving them, and one classic study of formal rater error training found it created new response sets and decreased accuracy. What works is structural. Use the same questions in the same order for every candidate. Define the criteria before you start. Score one dimension at a time rather than forming an overall view first. Use more than one rater. Where possible, remove the irrelevant cue from the process entirely, which is why blind marking works and why one study comparing in person courses with online courses found the appearance effect largely disappears when the instructor cannot see the student. The principle is to give the bias less material, not to ask people to be less biased.