Qoyod
Pricing
Qoyod
Pricing

Contrast Error in Appraisal

Term in Qoyod's Business Glossary. Practical definition with examples from the Saudi market.

What contrast error in appraisal is

Contrast error in appraisal is an appraiser measuring an employee’s performance against the employee they rated immediately before, instead of against the written standard for the rating.

Its mechanism is a moving reference point. The standard is supposed to sit outside the files and stay still; in practice it becomes mobile and shifts with every file that closes. An average employee who follows an outstanding one looks weak, and the same employee following a poor one would have looked strong.

Its distinguishing mark is that the rating changes when the order changes, while nothing changes in the performance, in the form, or in the appraiser. That mark is unusual among appraisal errors in being directly testable, and most of this page follows from the test.

The test: hold the file still and move where it sits

The nearest way to measure the effect is to fix the file and vary its position, because the effect is attached to position and to nothing else. The apparatus already exists in organisations that hold calibration sessions, since written example files are put in front of several appraisers at once in those rooms.

The figures below are assumed. Take a written benchmark file for an employee of average performance, where the organisation has agreed that the correct rating is 3 out of 5. Put that same file into two sessions, each with ten appraisers:

  • First session. It is placed second, after a strong file rated 5. The mean of the ratings given by those ten appraisers is 2.6.
  • Second session. It is placed second, after a weak file rated 1. The mean of the ratings given by those ten appraisers is 3.5.

The difference is 3.5 minus 2.6, which is 0.9 of a point. The scale runs from 1 to 5, so the distance between its ends is 4 points, and the difference is therefore 0.9 divided by 4, or 22.5 percent of the distance the scale covers. Both figures are in points on the same scale, which is what makes the division meaningful. Not one character changed in the file between the two sessions, nor in the form, nor in the descriptions of the levels. The only thing that changed was the file closed before it, and that is the only thing available to explain the result.

That percentage is not a threshold anybody should measure against. It is the size of the effect among those ten appraisers in those two sessions. Its value is that it moves the matter out of accusation and into measurement: the appraiser is not asked about their intentions, they are shown that the number moved while the file stood still.

A cheaper check when no benchmark file exists

Compare the mean rating of the first half of a session with the mean of the second half, then run the session again in the following cycle with the order reversed.

  • A session of ten files ordered by seniority: the mean of the first five is 3.8 and the mean of the last five is 2.9, a difference of 0.9, which on a four point span is 22.5 percent.
  • The following cycle with the order reversed: the mean of the first five is 3.6 and the mean of the last five is 3.1, a difference of 0.5, or 12.5 percent.

The two checks above were built on the same assumed gap, so the percentage they share is an artefact of the example and not a finding about anything.

What matters is not the size of the difference. It is that in both cycles the difference ran in favour of the first half whoever was in it. If the gap were an accurate description of particular people’s performance, it would have travelled with them when they moved to the end of the session. Its staying attached to the position is what makes the order, rather than the performance, the thing that explains the ratings.

This check works at the level of a session and not at the level of an employee, and it cannot be used to adjust one person’s rating. Its purpose is to tell the appraiser that their sessions need a break between files or a different order, not to raise the rating of whoever happened to come last on the list.

What the test would show for the errors nearby

The same experiment is the cleanest way to locate this error among the ones it is listed beside, because each of them would behave differently under it.

Reshuffle the files and central tendency error is unmoved: an appraiser who confines everyone to the middle does so in any order, and their ratings stay clustered wherever the files fall. That error measures against the right standard and uses only the middle of the scale, so it narrows the differences; this one widens them at times and inverts them at others, because it does not gather ratings anywhere, it attaches each one to its neighbour.

Lengthen or shorten the period and the recency effect in appraisal moves, while this error does not. That one takes its evidence from the wrong part of the period about the right person; this one may take its evidence from the whole period and then compare it with the wrong person.

And where an early impression about the employee drives the rating, the reference is a prior judgement about this person, formed before the evidence. The reference in contrast error is a different person entirely, who has nothing to do with the file that is open. A single impression carrying across every item on one person’s form runs in the other direction again: that movement is from item to item inside one person, while this one moves from person to person on a single item.

Where it appears most

  • Consecutive appraisal sessions. A manager opening their whole team’s files in one sitting closes each form with the last one still present.
  • Interviews on the same day. Two candidates in one day, and the second is measured against the first. Behavioral interview questions change what the answers are made of, since they ask about situations that actually happened, but they do not by themselves make two candidates’ scores comparable. A structured interview does, since every question carries a written description of what counts as a weak, an adequate and a strong answer, and that description does not move when the candidate does.
  • An assessment centre. Here the neighbour is not the file closed a moment ago but the person sitting across the table. In the group discussion several candidates work on one problem together while assessors watch, so the comparison is built into that exercise rather than smuggled into it. The arrangement carries its own answer, since several assessors observe against attributes described in behaviour that can be seen, recording independently before any of them speaks. It stops being an answer the moment the assessors discuss a candidate before writing anything down, which returns the room to a single moving reference.
  • A team with one visible star. That person becomes the ruler by default, and acceptable performance around them reads as falling short.
  • A team just after a weak performer leaves. Expectations rise without a decision, and the ratings of those who were appraised alongside that person fall.

Why the appraiser does not notice

Because comparing what is in front of you is a natural way to judge, and it does not feel like an error. The manager does not think that they lowered a rating because the previous person was better. They find themselves convinced that this employee is not at the level. The conviction is sincere, and the ruler it was formed against is the thing that moved.

So being warned about it is not enough, which is true of systematic rater leanings generally: what reduces them is a change in the instrument and the procedure rather than a change in anyone’s intentions.

Mechanics that hold the reference still

  • Rating one item at a time across the team rather than one whole file at a time, so that a single competency is rated for everyone against its written description and the answers are compared with the standard rather than with the last file.
  • Reading the level description again before each file, not once at the start of the session.
  • Separating consecutive interviews in time, and recording each candidate’s rating before meeting the next one.
  • Randomising the order of the files instead of ordering them by seniority or by expected rating, since a systematic order accumulates the effect in one direction.
  • Rating independently before consulting, because an opinion said first becomes the reference for everyone who speaks after it.

The room where this is most often caught is a performance calibration session, which puts shared examples in front of several appraisers together and so exposes a rating that nothing but its neighbour supports. That page’s own caution applies here in full: the examination starts with a question rather than a correction, and a difference between two files is not by itself this error.

Where the cost lands outside the appraisal form

  • A day of interviews. The candidate who follows an outstanding one leaves the shortlist while possibly meeting the written standard. This cost is never seen at all, because whoever was excluded does not come back to be compared again.
  • A promotion panel. The order in which candidates are presented becomes a factor in the outcome, and it is a factor that appears in none of the promotion criteria.
  • A decision at the end of a probation period. Where several cases are settled in one sitting, a person’s position in the running order enters a decision that in practice has no appeal.
  • A pay review. A rating that moved because of its neighbour passes into a table that builds a permanent difference on it, so the effect of a single session persists for years.

What resembles it and is not it

A standard that rose by a written decision moves the reference too, and it is the opposite case: if the organisation revised the description of a level and ratings fell, the reference moved by decision rather than by proximity. The difference is that this movement is written, dated and applied to everyone, while the contrast effect is unwritten and applies to whoever happened to be in a particular position.

A real difference between two consecutive files is ordinary. An outstanding performer and a weak one meeting in one running order is expected, and the second rating may be an accurate description. What establishes the error is a pattern that follows the order, not a single instance.

An instrument that asks for comparison explicitly is a different matter again. Where the task is to rank people against each other, comparison is precisely what was requested, and any objection to it is an objection of another kind. The error this page describes is a comparison occurring where the reference was supposed to be a written description of the rating.

And a declared external reference is not a neighbour at all. Measuring performance against a job level described in advance is measurement against something fixed and known before the first file is opened, which is the property the neighbour lacks.

What we did not find

We do not lay down here a numerical point above which a difference becomes evidence of this error, nor a number of files at which a session becomes too long, nor how much separation between two interviews is enough. We did not find, in our sources, anything establishing any of those, and the figures used above are example figures showing how the calculation is done rather than thresholds organisations should be measured against.

What the page does say is that the effect can be measured by holding the file still and moving where it sits, and that reading the result remains an administrative judgement made inside the organisation.

Qoyod HR

A standalone Saudi HR system

One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.

Explore Qoyod HR

A standalone system on its own subscription. The connection to Qoyod Accounting is now available.

Related terms

Ready to apply accounting the right way?

Qoyod runs your accounting with precision and full ZATCA compliance

Try Qoyod free for 14 days — No credit card required.