What peer appraisal is
Peer appraisal is asking an employee’s colleagues at the same level about their conduct in the work they share, so that those answers enter the employee’s performance record alongside their direct manager’s rating.
It is one instrument among those that widen the angle of observation, and what distinguishes it is that its source has no authority over the employee and holds no decision about their pay or their promotion. That single fact is both where its strength comes from and where everything that spoils it comes from.
Why a colleague is asked at all
The argument for this instrument is an argument from observation, not an argument from fairness. A direct manager mostly sees outputs: what was delivered, when it was delivered, and whether it was accepted. A colleague who shares the daily work sees the route to the output: how information was asked for, whether there was an answer when one was needed, and whether what had to be handed over was handed over before the leave or after the task closed.
So the question a colleague answers is not whether the person performed. It is how they performed alongside the people around them. An organisation that takes the manager’s rating alone judges the employee from a single distance, and that distance shows the product while hiding what it cost the team.
It is worth being clear that widening the angle in this direction is not the same as widening it upward. Asking the people a manager supervises about that manager runs up the line of authority, where the rater is weaker than the person rated, and it produces a record about a different subject entirely. That instrument is subordinate appraisal, and the pressure on it is a different pressure: fear of a manager is not the same thing as courtesy between equals, and the two are not treated alike.
The first decision is how many people are asked
The effort usually goes into wording the questions, while the number that decides how stable the result is happens to be how many people are asked. Take an employee rated by colleagues on a five point scale, where four of them agree on 4 and one dissents at 1:
- With five raters, the ratings sum to 17 and the mean is 3.4.
- With three raters, two at 4 and one at 1, the mean is 3.0.
Against a baseline in which everyone had agreed at 4, the single dissenter pulled the mean down by 1.0 of a point when there were three raters, and by 0.6 of a point when there were five. The rule behind both figures is simple: the effect of one dissenter is the size of their deviation divided by the number of raters. A deviation of 3 points reduces the mean by 1.0 with three raters, by 0.6 with five, and by 0.5 with six. The denominator is the count of people who answered, not the count of people who were invited.
What follows is that raising the number of raters buys stability, and that what is bought shrinks as the number rises: the distance from three to five is larger than the distance from five to six. Anyone wanting more stability past a certain point is paying for it in the time of many colleagues in exchange for a small movement in a number.
The minimum is an administrative choice, and its reason is concealment
A minimum number of raters below which the result is not shown is common in this area, and it is the organisation’s own choice, set according to the size of its teams. It is not a settled rule we are relaying from a source, and it is said here before any number rather than after one, because a number written first becomes a fixed threshold in the reader’s mind whatever qualification follows it.
The reason that minimum exists is concealment rather than arithmetic. An employee who knows that three people were asked about them, and who sees one unusual rating in the result, can work out whose it was by elimination. Once that inference is available, candour ends in the following cycle, and what gets collected afterwards is a set of close and courteous ratings whose effect shows up as central tendency error.
What a colleague can be asked, and what they cannot
What fits is what falls inside what they see: cooperation when asked, transfer of knowledge, keeping what was promised to a colleague, and their effect in a team meeting. What does not fit is what they do not see: the size of anyone’s contribution to a financial result, their capacity for promotion, and the quality of technical work outside their own field.
A form that mixes the two produces numbers that are meaningless in half of it, and then averages that half with the other half into a single figure that looks like a sound number. So the question is tied to a described behaviour, which is the same condition a graphic rating scale has to satisfy when its levels are written as statements rather than as bare numbers, and the same condition the organisation’s competencies have to satisfy to be usable at all.
The mean hides the spread, and the spread is the news
Take two employees, each rated by five colleagues. The first is rated 4, 4, 4, 4 and 1. The second is rated 3, 3, 4, 4 and 3. The ratings sum to 17 in both cases and the mean is 3.4 in both cases.
The two situations are nevertheless opposites. The first is agreed on except by one person who sees something the others do not, and the distance between the highest and lowest rating is 3 points. The second is agreed within a narrow band, and its distance is 1 point.
The decision taken in the two cases is not the same either. The first calls for a question about where the disagreement comes from; the second calls for nothing. So whoever carries only the mean into the performance record has dropped the news and kept the number. What travels with every mean is the number of raters and the distance between the highest and lowest rating, and those are two lines rather than a redesign.
The cost is time, and it is a computable amount
Take a department of 40 employees, where it has been decided that each is rated by 5 colleagues. That is 40 multiplied by 5, or 200 forms. If one form takes 12 minutes, the total is 200 multiplied by 12, or 2,400 minutes, which is 40 hours of working time in a single cycle. The forty hours and the forty employees are different quantities that happen to coincide in this example, and neither is derived from the other.
Those hours come out of the employees’ own time rather than out of the HR function’s, which is the part that gets overlooked when the cycle is designed. Two cycles a year double it. And raising the number of raters from five to six makes 240 forms, so the 40 extra forms cost 480 minutes, or 8 further hours, in exchange for the movement from 0.6 to 0.5 in the earlier example, which is 0.1 of a point of stability.
So the number of raters and the number of cycles are decided together rather than separately, and what is bought is set against what is paid. A cycle designed at its fullest, for which colleagues then have no time, gets completed in two minutes a form, and the stability that was purchased with the number of raters is lost again in the haste of the answers.
Where it does not hold up
- Work that is not shared. An employee working alone in a branch, or on an assignment where they meet nobody, gives their colleagues nothing to have observed, so what is collected about them is an impression rather than an appraisal.
- A very small team. Where the number of colleagues does not reach the point at which the person answering is concealed, the instrument is running with no concealment available, and a direct conversation is the better route.
- A year in which the teams were restructured. The colleagues of the first half are not the colleagues of the second, so the raters are describing two different periods and the descriptions are pooled into one number.
- A first use tied to a pay decision. The first cycle is what teaches people the instrument, and an organisation that attaches money to it from the outset has taught them to be careful with it before they have learned to use it.
Instruments that leave a different record behind
The quickest way to separate this from what sits beside it is to ask what each one deposits in writing when it has finished running.
A peer review panel deposits a recommendation about a decision already taken in one worker’s case. It runs because an event occurred, and it produces a view on that event. Peer appraisal runs at a fixed date whatever did or did not occur, and it deposits a periodic rating of conduct.
Continuous feedback between two colleagues deposits nothing into a total at all: it happens between them at the time and stays there. Peer appraisal is collected, computed and entered into a record. The two can both be running in the same organisation, and neither substitutes for the other.
What spoils it
- Reciprocity. Two colleagues who each know they are being rated by the other drift toward a mutually high rating. This is the heaviest thing that happens to the instrument, and the remedy is that the person rated does not know who was asked about them.
- Coupling it to fixed proportions. Where peer results are attached to a quota to be filled, the colleague becomes an opponent in a division, and lowering their rating becomes a direct interest of whoever is rating them.
- Short memory. A colleague is asked once a year about conduct that lasted a year, so their answer is dominated by the recent weeks, which is the recency effect in appraisal exactly.
- Letting the person rated choose their raters. Whoever lets an employee name who rates them has collected the opinion of their friends at work, which is a genuine opinion and is not a sample of their team.
- Pooling without calibration. A generous team and a severe one pooled into one table show a difference between teams as performance when it is a difference in rating style, which is what a performance calibration session is for.
Before the result goes into the performance record
Four things are written down with it: how many raters were asked, who chose them, the distance between the highest and lowest rating, and what behaviour the question was about. Without the first, the stability of the number is unknown. Without the second, the sample is unknown. Without the third, agreement is read where disagreement exists. Without the fourth, the number is read as a judgement on a person rather than as a description of a behaviour.
Then, before the first cycle rather than after it, the organisation decides what of the result is shown to the employee and what is not. An organisation that defers this until the answers are in finds itself between two bad options: withholding what it promised to disclose, or disclosing what identifies whoever said it. Either one ends the instrument in a single cycle.
What is left open
We did not find, in our sources, a published minimum number of raters for the Saudi market, and what is set out above is a calculation of an effect rather than a recommendation of a figure. Whether peer results enter a pay decision is a policy question inside the organisation, and it carries a price in either direction.
The numbers in the worked examples are assumed in order to show the method. They are not a standard, and they are not taken from any source.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.