Qoyod
Pricing
Qoyod
Pricing

Central Tendency Error

Term in Qoyod's Business Glossary. Practical definition with examples from the Saudi market.

What central tendency error is

Central tendency error is an appraiser confining their ratings to the middle of the scale, so that everyone on the team receives an average rating or something close to it however differently they actually performed.

Its mechanism is range compression. The scale offers five levels and only the second and third get used. No individual rating has moved to a wrong position; what the ratings have lost is the ability to separate one person from another. The fault is in how much of the scale was used, not in where any one rating sits on it.

Compression and shift are different faults

This error is routinely confused with an appraiser’s tendency toward leniency or severity, and the difference is in the kind of fault rather than its size. Compression pulls the ratings together, and the place they gather is the middle; what is lost is the difference between people. Leniency and severity shift the ratings, all of them upward or all of them downward, and the spread between them can stay wide; what is lost is where the team sits on the scale, not how its members compare with each other.

Separating the two has a practical consequence. A severe distribution keeps its ability to order people, so it still works as an input to a decision that rests on comparison, even while it describes the team unfairly. A compressed distribution has lost the ability to order people at all, so it works as an input to nothing, wherever on the scale it happens to sit. The two are therefore not treated alike: a shift is corrected against an external reference the appraiser can be compared to, and compression is corrected by widening what the appraiser is able to support with evidence.

The same criterion separates it from the errors it is usually listed beside. A pattern that would change if the files were reshuffled is about the order of the files, which is contrast error in appraisal, and compression does not change when the order does: an appraiser who rates everyone in the middle does so in any order. A pattern that would change if the appraisal period were lengthened or shortened is about the window the evidence came from, which is the recency effect in appraisal, and compression is indifferent to the length of the period. And where a single impression carries across all the items on one person’s form, the movement is between items within one person; compression moves between people on one item.

How it is recognised

Not from a single rating. An employee of average performance deserves the average rating, and giving it to them is not an error. What reveals the pattern is the shape of the distribution: a whole team’s ratings piled on one level or two adjacent ones, both ends of the scale empty year after year, and a wider range from another appraiser working on the same form in a comparable set of roles.

That last comparison is what makes it an observation rather than an accusation, and it is one of the things a performance calibration session exists to surface, since a whole team clustered in the middle of the scale is visible in that room and invisible from inside the appraiser’s own forms.

Why it happens

  • The two ends cost more than the middle. A high rating gets asked for its justification, a low one opens a difficult conversation, and the middle is asked for nothing.
  • There are no recorded observations. An appraiser who wrote nothing down during the year has nothing to support a rating at either end with, so they retreat to the one that needs no support.
  • The scale carries no descriptions. A ladder of bare numbers with no behavioural description at each level makes the middle the default answer, which is why the design of a graphic rating scale is where this error is usually admitted rather than where it is committed.
  • One manager rates too many people. A wide span of control makes distinguishing between individuals expensive in time alone, before any question of courage arises.

What a raise decision loses when the range narrows

The cost does not appear on the form. It appears in the table built from the form, and the figures below are assumed in order to show where.

Take a team of eight, a scale of five levels, and an annual raise allocation of SAR 40,000 distributed in proportion to the ratings. Each person’s share is the allocation multiplied by their own rating and divided by the sum of the team’s eight ratings, so the denominator throughout is a total of rating points and not a headcount.

First case, with the range compressed. The appraiser gives seven of the eight a 3 and one of them a 4, so the eight ratings sum to 25. The share of a person rated 3 is 40,000 x 3 divided by 25, which is SAR 4,800. The share of the person rated 4 is 40,000 x 4 divided by 25, which is SAR 6,400. The gap between the top of the team and everyone else is SAR 1,600, and the highest share is one and a third times the lowest.

Second case, same team and same allocation. The appraiser uses the full scale, and the eight ratings come out as 5, 4, 4, 3, 3, 2, 2 and 1, summing to 24. The share of the person rated 5 is 40,000 x 5 divided by 24, which is SAR 8,333 to the nearest riyal. The share of the person rated 1 is 40,000 x 1 divided by 24, or SAR 1,667. The distance between the two ends is SAR 6,667, and the highest share is five times the lowest.

Nothing changed between the two cases except how much of the scale was used. Not the allocation, not the size of the team, not the form. The same money left the first case distributed in a way that is barely visible: not wide enough to read as a distinction, and not absent enough to read as a declared equality. How much of a rating difference becomes a pay difference is set by the organisation’s own policy and not by the scale, which is a separate decision taken in the design of the salary structure. What the compression decides is whether there is a difference for that policy to act on.

Measuring the compression before it reaches the money

The pressure on the instrument can be read off the ratings table on its own. A scale of five levels makes five values available. If only two adjacent levels are used, then two of the five available levels were used, which is 40 percent of them; if three are used, 60 percent. The figure is computed directly from the ratings and needs no judgement about any particular appraiser, which is what makes it a reason to open a question rather than an answer that closes one.

It is worth being exact about what that percentage counts, because a second and different figure lives in the same table. Counting levels used against levels available is not the same as measuring the distance covered: two adjacent levels on a scale of 1 to 5 span one step out of the four steps that separate the ends. Both are true, they answer different questions, and quoting one while naming the other is how a defensible figure turns into a misleading one.

When clustered ratings are not this error

The question that decides it is whether the pattern belongs to the appraiser or to what is being appraised, and there is a direct way to ask it: does another appraiser, working on the same form across a comparable set of roles, produce a wider range?

Several familiar situations fail that question before it is even reached, because they do not supply a pattern to test. A team of three or four may genuinely be close in performance, and their ratings clustering on two levels carries far less than the same clustering in a team of twenty; the smaller the group, the less a distribution can establish. An appraiser in their first cycle, who took over a team two months before it began, has fewer observations than the period requires, and thin evidence is not a compressed range. That is a problem about evidence before it is a problem about the scale, and its proper home is the window those observations came from. A small unit selected on narrow criteria, from which those who stayed have stayed, may have close ratings as an accurate description. And a single form establishes nothing in either direction, since an average performer deserves the average rating.

Where the comparison can be made, it usually settles the matter. If the other appraiser’s range is wider across a similar team, the narrowness is a property of the appraiser. But the case worth planning for is the one where both ranges are narrow, because it passes the test and still does not point at either appraiser. The question then moves up a level, from the appraiser to the instrument, and from the instrument to what the organisation actually demands by way of justification at each end of the scale. An organisation that asks for a written case at the top of the scale and for nothing in the middle has built the incentive, and it will keep producing the same distribution with any appraiser it hires.

Widening what an appraiser can support

A behavioural description written for every level turns choosing a level into matching a description rather than estimating a number. A running record of observations kept through the period gives an appraiser the evidence to defend both ends. Rating one item at a time across the whole team, rather than one person’s whole form at a time, puts the differences between people where the comparison naturally belongs. Asking for an example at every rating, rather than only at the high ones, removes the asymmetry that made the middle cheap. And training managers in the difficult conversation addresses the commonest cause directly, since avoidance of that conversation, rather than any weakness in the instrument, is what most often produces the middle.

Treating it by imposing proportions on the appraiser swaps one error for another. Forcing distributions to match until they align produces an apparent fairness that penalises the team that genuinely performed better. What is needed is to widen what the appraiser can support with evidence, not to oblige them to produce a set number at each level.

Where this page stops

We do not lay down here a correct shape for a team’s distribution of ratings, nor a proportion each level is supposed to reach, nor a threshold below which clustering is acceptable and above which it is an error. We did not find, in our sources, anything establishing such a figure for the private sector in the Kingdom, so we stop at what we have: a description of the mechanism, a way of reading the pattern, and what narrows it.

Any figure offered as the natural distribution of a team is an administrative choice at bottom, and it should be said as a choice. That imposing such proportions is a remedy which swaps one error for another has already been set out above, and this refusal is its other face.

Qoyod HR

A standalone Saudi HR system

One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.

Explore Qoyod HR

A standalone system on its own subscription. The connection to Qoyod Accounting is now available.

Related terms

Ready to apply accounting the right way?

Qoyod runs your accounting with precision and full ZATCA compliance

Try Qoyod free for 14 days — No credit card required.