Qoyod
Pricing
Qoyod
Pricing

Performance Calibration

Term in Qoyod's Business Glossary. Practical definition with examples from the Saudi market.

What performance calibration is

Performance calibration is a session in which the appraisers meet before performance ratings are confirmed, agree through shared examples what each rating means, and examine the differences between their distributions while the ratings can still change.

Everything about the session follows from its position in the cycle. It sits after the forms are filled in and before anything is announced, and that is the only window in which its work is possible.

The problem it was built for

One form does not produce one ruler. Two managers using the same form can mean two different levels of performance by the same rating, and nothing on the paper reveals it. The disagreement only surfaces when the outputs of their two teams are placed side by side, at a promotion, a development route, or any decision built on the rating.

At that moment the difference between two employees has become a difference between their managers. Neither manager did anything careless. They simply read the same scale differently for a year, in separate rooms, with no occasion on which the readings met.

The session addresses this before it happens rather than after. Revising ratings once they have been given to employees is a late correction, and it is read as a retreat from something already promised. The same act, performed a week earlier, is ordinary quality work on an unfinished process.

What happens before the session

The preparation is what decides whether the meeting is possible at all. Each appraiser arrives with written examples rather than with their recollection, because the session compares reasons and a recollection cannot be compared with anything.

The scope is set in advance and kept narrow enough to be meaningful. Comparable jobs are calibrated together; setting roles of different natures against each other produces a comparison with no content, and it makes any difference that appears impossible to interpret. The items being compared also have to be the same items, which is where the design of the form itself constrains what the session can do: a graphic rating scale whose levels carry no descriptions gives the room nothing to converge on, because there is no written statement of a level for two people to agree or disagree about.

Someone chairs it who has no team in the room. The chair’s job is to keep the question falling on everyone equally, and an appraiser who is also defending their own distribution cannot do that while doing the other thing.

What the session settles, and what it is not allowed to settle

The subject of the session is the ruler: what above expectations means in this organisation, and by what evidence it is recognised. An example employee is presented at each level and the room discusses why that person was placed there, until the standard is understood through cases rather than through the wording of a description.

A rating that changes afterwards changes as a consequence of correcting how the standard was applied to a particular case. It is not the purpose of the session and it is not a quota to be reached. A session that opens with a number of ratings to be moved has already decided the thing it was convened to examine.

This is the whole distance between calibration and a forced distribution, and the distance is worth stating plainly because the two are run in the same room with the same people. A forced distribution fixes in advance what proportion of people fall in each rating, which obliges the appraiser to produce a known count whatever their team did. Calibration imposes no proportion at all. It asks what explains a difference once a difference appears.

Confusing the two is what turns the session harmful rather than merely unproductive. Levelling distributions by force until they match produces an apparent fairness that penalises the team that genuinely performed better, and it does so under the authority of a process designed to protect against exactly that. A difference between two appraisers is not by itself evidence of a distorted ruler: one team may be stronger, or its work harder. The examination begins with a question and not with a correction, and an organisation that cannot tell those two openings apart should not convene the session.

Two further rules keep the room honest. The reason for every change is written down, because a rating that moved with no recorded reason cannot be explained to the person it belongs to. And each case carries a time limit, or the first cases consume the entire session and the last ones are approved by exhaustion.

What the room can see, and where its sight ends

Comparing appraisers’ averages against each other is what makes a consistently lenient or consistently strict rater visible, and neither is visible from inside their own forms. A whole team’s ratings piled into the middle of the scale is visible in the same way, which is how the session surfaces central tendency error, and the entry on contrast error in appraisal puts the room to a related use, since shared examples shown to several appraisers at once expose a rating that only its neighbour in the running order accounts for. A standard that nobody actually understands announces itself the moment the appraisers disagree about a single example, which is a more useful signal than any of the distributions.

The boundary is in the material, not in the method. The session works on what the appraiser recorded during the year, so it reaches everything that was written down and nothing that was not. An appraiser who observed little and recorded less brings little to the table, and no amount of discussion in the room manufactures the evidence that was never gathered. The session also does not remove the systematic leanings any human rater brings. What it does is put their effect where other people can ask about it, which is a smaller claim than removing them and a more defensible one.

What voids the session

The commonest way to void it is to run it by seniority in the room. Whoever holds the higher post or the louder voice wins the room’s opinion, and the rating moves from the bias of one appraiser to a collective bias that looks more legitimate because a group produced it. Nothing in the session’s design prevents this on its own; only the chair and the written reason for each change do.

The second way is to detach it from the employee. Calibration is an internal procedure between appraisers, and what reaches the employee is their rating and its reasons, from their own manager. If the rating arrives altered with no explanation, the session has corrected the ruler and destroyed confidence in the result, which is a net loss even when the new rating is the more accurate one.

What comes out of it besides the ratings

The output is not the confirmed scores alone. Whatever disagreement about meaning surfaced in the room goes back into the appraisal standards for the following year, which is the point at which a session stops being an approval stage and becomes a source of information about the appraisal system itself. That is also the route by which the room’s findings reach the written performance standards, since a level that two appraisers read differently is a level whose description was not doing its work.

Repeated performance gaps concentrated in one team are read the same way, as an organisational cause rather than an individual one. A room that keeps finding the same shape in the same unit is being told something about that unit, and the ratings are the least interesting part of the message.

After the session the rating remains an input to later decisions rather than a decision: a development route, a formal improvement process where the gap has been described, or a difference in pay that needs a documented reason somebody can say out loud. That last one is where the session pays for itself, because a difference in the salary structure resting on a rating nobody can explain is a difference that will eventually have to be explained.

The same logic reaches ratings collected from more than one source. The entry on peer appraisal makes the point from its own side: pooling the peer results of a generous team and a severe one into a single table, without this step in between, turns a difference in rating style into an apparent difference in performance.

What is not settled here

We did not find, in our sources, a required form for such a session, nor a prescribed frequency, nor any rule obliging an organisation to hold one. What is described above is practice as it is observed, not a rendering of any published instrument.

Nor does this page set out a correct shape for a team’s distribution of ratings, or a point at which a difference between two appraisers becomes evidence of anything. Any figure presented as the natural distribution of a team is an administrative choice at bottom, and it should be stated as a choice by whoever adopts it.

Qoyod HR

A standalone Saudi HR system

One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.

Explore Qoyod HR

A standalone system on its own subscription. The connection to Qoyod Accounting is now available.

Related terms

Ready to apply accounting the right way?

Qoyod runs your accounting with precision and full ZATCA compliance

Try Qoyod free for 14 days — No credit card required.