Qoyod
Pricing
Qoyod
Pricing

Graphic Rating Scale

Term in Qoyod's Business Glossary. Practical definition with examples from the Saudi market.

What a graphic rating scale is

A graphic rating scale is an appraisal form on which the items being rated are listed, and each one is marked at a single position on a fixed ordered scale: from below expectations to above expectations, for instance, or on a numbered scale from 1 to 5.

It is called graphic because the original form was a line or a row of boxes with one position marked on it, so the whole sheet could be read at a glance. It is among the oldest appraisal form designs and among the most widely used, for the same reason in both cases: it is quick to fill in and quick to total.

What the form is made of

  • The items. What is being rated. These are competencies, or responsibilities taken from the job description, or a mixture.
  • The scale. How many levels there are and how they are ordered, and it is the same scale for every item on one form.
  • The level descriptions. What each level means. This is the part most often left out, and whether the form works at all turns on it.
  • The weights, where there are any. Whether the items count equally toward the final score.
  • The evidence field. Written space in which the choice is supported. It is optional on most forms and necessary on the useful ones.

Two of those five carry most of the weight in practice. Level descriptions have to be written per level and per item, not per level alone: “good” without a description means two things to two managers, and it means them without either manager noticing the disagreement. Items have to come from what the job requires, and the test for a usable item is the one the entry on competencies sets out for a usable competency, which is an observable behaviour at a stated level of proficiency with a connection to the job. An item phrased as an adjective, such as diligent or cooperative, cannot be observed, so it gets filled in from a general impression of the person rather than from anything they did.

What spoils the form in use is mostly the mirror of that. A scale stripped of its descriptions leaves each rater to invent their own ruler, and it makes the middle of the scale the option nobody has to justify. A form of thirty items filled in over a few minutes produces one repeated column of choices. Totals summed from items of unequal importance hide the differences the form was built to show. And filling in the form one person at a time, rather than one item at a time across the team, lets the ruler drift between files.

What the weights do to a score that never changes

The arithmetic below is assumed, and it is worth running because the result is not visible from the form. Take a five item form on a scale of 1 to 5, with weights of 30, 25, 20, 15 and 10 percent, and one employee rated in that order at 4, 3, 5, 2 and 3.

The weighted score is (30 x 4 + 25 x 3 + 20 x 5 + 15 x 2 + 10 x 3) divided by 100, which is (120 + 75 + 100 + 30 + 30) divided by 100, which is 355 divided by 100, or 3.55.

The simple mean of the same five ratings, if the weights are ignored, is (4 + 3 + 5 + 2 + 3) divided by 5, which is 17 divided by 5, or 3.40. The two methods differ by 0.15 of a point on identical ratings.

Now reverse the weights, so they run 10, 15, 20, 25 and 30 percent, without changing a single rating. The score becomes (40 + 45 + 100 + 50 + 90) divided by 100, or 3.25. The distance between 3.55 and 3.25 came from neither the employee’s work nor the rater’s judgement. It came from the order of the weights alone.

The ceiling on what any one item can move

The more useful thing the same table yields is a limit nobody reads off it. The most an item can move the final score is its weight multiplied by the distance between the ends of the scale. On a scale of 1 to 5 that distance is four points, so an item weighted at 10 percent cannot move the final score by more than 4 x 0.10, which is 0.40 of a point, even if it travels from the bottom of the scale to the top. An item weighted at 30 percent can move it by up to 4 x 0.30, which is 1.20, or three times as far.

What follows is about where the appraisal meeting spends its time. An item can absorb a long discussion while being arithmetically incapable of changing the outcome, and another item can be filled in within seconds while carrying three times the effect. Both facts are computable from the weights table before a single form is filled in, which makes the weights an agenda decision rather than an administrative one.

When the score becomes a band, one item decides the band

Most organisations do not use the number as it stands. They convert it into a band: below expectations, meets expectations, above expectations. Suppose the threshold for the top band sits at 3.60. The score of 3.55 does not reach it, and the distance from the threshold is 0.05 of a point.

Raise the second item alone from 3 to 4, at its weight of 25 percent, and that adds 25 x 1 divided by 100, which is 0.25. The score becomes 3.80 and crosses the threshold. One level on one item moved the employee from one band to the next, and with it moved whatever the organisation attaches to the band.

This is where rounding starts to matter. The distance between 3.55 and 3.60 means nothing as a description of performance and means everything in the table that reads the band. So the thresholds belong in the description of the form, along with who may change them and when, because moving a threshold by a tenth of a point reclassifies a whole team without a single form being altered.

The form is where the score is written down, not where it comes from

A graphic rating scale is a container. What it gets asked to supply mostly belongs to instruments around it, and the quickest way to see which is to ask what part of the appraisal each one owns.

The score has a source, and the source is an act of observation. A 360 degree appraisal widens who does the observing and can collect each source’s answers on a graphic rating scale, so the two sit on top of each other rather than competing. A self assessment is an input from the employee, better used to open a conversation than to enter an arithmetic total.

The score also has a procedure that steadies it, and that procedure works on what comes out of the forms rather than inside them. Performance calibration takes the completed forms and settles what a level means in this organisation before the scores are confirmed. This is why replacing one form design with another does not fix raters disagreeing about the ruler: the disagreement is in how the scale is read, not in how it is drawn. The same point explains why a scale without level descriptions is the entry route for central tendency error, since a rater with no described level to match has only a number to estimate, and the middle of a numbered scale is the estimate that costs nothing to defend.

And the score has a content decision behind it, which is what the items are. A competency model decides what gets measured and in what terms; the form decides how what was measured gets recorded. Swapping the container leaves the content where it was. A set of agreed objectives is content of a different kind, settled at the start of the period and scored by whether it was reached, while the form is filled in at the end. A checklist of statements marked yes or no has no ordered scale and no levels at all, so what it produces is a count of what applied rather than a position on a line. A ranking instrument that orders people against each other, or distributes them into fixed proportions, is answering a question about the team; the graphic rating scale rates each person against the scale alone and does not know who else is in the team.

One neighbour looks nearer than it is. A satisfaction or self rating scale in a survey has the same shape and a different subject: it records a stated opinion of the person giving it, while this form records one person’s judgement of another person’s work and carries consequences. The shape is shared and nothing else is.

The filled form is the record, and it is an input

The completed form is what remains in writing from the entire cycle, which means it will be read a year later as the official description of that period’s performance. That is a load it was not designed for if the evidence fields are empty: a grid of ratings with no examples attests that an appraisal happened and attests to nothing that occurred. The design principle that follows is the one on the item side as well, since a form that accepts a choice without an example accepts an impression. That is also what makes a structured interview and a well built appraisal form the same kind of object: in both, the described level is what the answer is compared against, and it is fixed before the person arrives.

Its other limit is that it is an input and not an output. What follows from weak performance, whether a development route or something more formal, has its own written procedure and its own conditions outside the form, and the rating does not stand in for them. The rating says there is something to look at. It does not say what is to be done or by what route.

What is not claimed about the instrument

We did not find, in our sources, a figure for how much validity or reliability this instrument carries, nor a numerical comparison against other form designs, nor a best number of levels for the scale. The literature that addresses those questions is outside our sources, so it is neither reproduced nor summarised here.

What is set out above describes what the instrument is, what it is used for, and what is observed when organisations use it. It is not a judgement about its measurement properties, and the figures in the worked examples are there to show the method rather than to serve as values anyone should adopt.

Qoyod HR

A standalone Saudi HR system

One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.

Explore Qoyod HR

A standalone system on its own subscription. The connection to Qoyod Accounting is now available.

Related terms

Ready to apply accounting the right way?

Qoyod runs your accounting with precision and full ZATCA compliance

Try Qoyod free for 14 days — No credit card required.