What test reliability is
Test reliability is the consistency of the score an instrument produces: that it returns a similar result for the same person under equivalent conditions, and that the ranking of candidates does not move for reasons that have nothing to do with the candidates.
Its question is stability and stability alone. It tells you the instrument gives the same reading. It tells you nothing about what is being read.
Start with what an unstable score does to a shortlist
Take six candidates for three seats, whose papers are marked independently by two markers working from the same key. The scores come out as follows.
- Ahmed. Marker A 82, Marker B 78, a gap of 4.
- Badr. Marker A 80, Marker B 84, a gap of 4.
- Khalid. Marker A 79, Marker B 71, a gap of 8.
- Saad. Marker A 77, Marker B 81, a gap of 4.
- Fahd. Marker A 76, Marker B 79, a gap of 3.
- Majed. Marker A 74, Marker B 72, a gap of 2.
Marker A’s top three are Ahmed, Badr and Khalid. Marker B’s top three are Badr, Saad and Fahd. Of the three names that would be shortlisted, one survives the change of marker. Two candidates enter the shortlist and two leave it, and neither the papers nor the key moved at all.
Now look at the size of the disagreement that produced it. Across the six candidates the largest gap is eight points, at Khalid, and the smallest is two, at Majed. The mean gap over those six candidates is 4.17 points, which is a number most people would describe as small. The gap separating the third seat from the fourth on Marker A’s list is two points. The instability did not have to be large. It only had to be larger than the margin the decision was being made on.
That is the operational rule this whole concept reduces to. Instability is not judged by its absolute size but by its size relative to the margin you are separating accept from reject with. An organisation choosing between candidates a point or two apart needs stability that an organisation dropping everybody below half and taking the rest simply does not.
The scale analogy, and the limit inside it
A scale that reads five kilograms for a box one time and seven the next is useless before anybody asks whether it is weighing the right thing. That is reliability: a property of the instrument that sits logically in front of any discussion of what it measures.
The analogy carries a limit worth stating, because it is the reason most organisations have no idea how stable their instruments are. A faulty scale does not reveal itself in a single reading. Anyone who weighs the box once and sees a plausible number holds nothing capable of exposing the variation. A test administered once to one cohort is in the same position: it has a result, and the result contains nothing that indicates whether it would recur.
Four forms, and how to tell which one applies to you
- Consistency over time. The same group sits the instrument again after an interval, and the two results agree. The interval has to be long enough that the questions are forgotten and short enough that the thing being measured has not itself changed. Too short and it measures recall of the answers; too long and it mixes instability with real change in the candidate.
- Equivalence of forms. Two versions with different questions from the same domain produce similar results. This is what you need when the instrument runs across successive intakes, or when a candidate sits it a second time. It is weaker in practice than it is assumed to be, because the second form is usually written in a hurry once the first has circulated.
- Internal consistency. The items within one instrument pull in one direction, so that no two of them diverge without a reason. This is the only one of the four that can be examined from a single administration, which makes it the accessible one for an organisation that cannot run the instrument twice.
- Agreement between markers. Two independent markers reach the same score for the same answer. This is the decisive form for anything not machine marked, including every structured interview scored against a written scale, and it is the only one of the four where the instability lives entirely outside the candidate’s paper.
All four are not required in every case. What decides which one to examine is the shape of the use. An instrument marked against an open key is asked about marker agreement before anything else. One running across successive intakes with refreshed questions is asked about equivalence of forms. One feeding a decision months after the sitting is asked about consistency over time. Examining a form that is not where your instrument’s problem lives produces a reassuring answer to a question you were not asking.
Stable score and stable ranking are two different properties
The table above contains two questions that are easy to run together. The first is whether the same person gets the same number. The second is whether the order of candidates holds even if the numbers move.
They can move together and either can move alone. A severe marker who takes five points off everybody breaks the first and leaves the second intact: the names stay in their places and the figures are lower. A marker who is severe on one kind of answer and not another breaks the second even while the average score is preserved exactly.
Which of the two matters to you follows from how the score is used. An organisation ranking candidates for a fixed number of seats cares about the order. An organisation applying a written threshold and admitting everybody above it cares about the level of the score itself, because a general downward shift of five points moves a whole group from above the line to below it.
What weakens it in practice
- Too few items. A three question instrument is at the mercy of a candidate’s luck on any one of them, because each question carries a third of the score.
- Ambiguous wording that each reader takes a different way, so the instrument measures comprehension of the question rather than command of its subject.
- An open marking key that leaves wide discretion with no worked examples to anchor it.
- Administration conditions that vary between one candidate and the next: a quiet room and a crowded one, a machine that works and one that fails.
- A scale with too few points, which puts candidates of genuinely different performance on the same number.
- Marker fatigue, where a long batch marked in one sitting is judged more strictly at the start than at the end.
- Order effects, where a mediocre paper read straight after a weak one looks better than it is.
Almost every item on that list lives in the design or the administration of the instrument rather than in the candidates. That is why a stability problem is addressed by reviewing the instrument before reviewing the people who sat it.
The remedies are procedural, not statistical
Very little of the fix is arithmetic. Adding items within the same domain reduces the share of the score any single item carries. Expanding the key with worked examples of a weak, an acceptable and a strong answer makes two markers read the same thing. Training markers on shared papers before live marking, and arguing about the disagreements then, moves the argument to the point where it costs nothing. Writing the procedure down, including the time allowed, the instructions read out and the materials permitted, takes it out of the hands of whoever happens to be running the room. Concealing the candidate’s name from the marker where that is possible removes one input that should never have been one. And splitting the score into named components rather than a single overall figure means the disagreement surfaces in a specific component and gets fixed there.
That last point is the one with the widest application, and it is the same discipline a work sample test depends on and that the wash up session in an assessment centre exists to exploit: a total tells you that two markers disagreed, and only the components tell you where.
Properties that get read off it and are not in it
Objectivity is a nearby property that is cheaper to obtain. A machine marked instrument is objective by construction, so the absence of a human marker buys consistency in the marking and leaves the consistency of the reading untouched: an instrument whose questions are few or ambiguous stays unstable with a machine reading it. Difficulty pulls in the other direction entirely, since an instrument that is very hard bunches everybody into a narrow low band, which weakens its ability to separate candidates while saying nothing about whether its reading repeats. And agreement between two different instruments about the same person is the one that looks closest and sits furthest away, because that is a question about what the two instruments measure and belongs to construct evidence.
High reliability does not make a test fit for use
An instrument can be perfectly stable while measuring something with no bearing on the job. A typing speed test returns a very consistent result, and it does not support a decision about a role in which nothing is typed.
Stability is a property of the instrument; the connection to the decision is the separate question that test validity deals with. Treating the first as an answer to the second is what produces well built instruments of no use, and it happens for an unremarkable reason: stability can be examined with numbers the organisation already holds, and what lies beyond it requires a judgement or a wait. So the examinable question gets examined, and its answer is filed against a question nobody asked.
When to look at it again
Stability is lost rather than acquired once, so it is revisited on four events: a change in who marks, a change in the medium from paper to screen or from a room to a remote sitting, an edit to the questions or the removal of some of them, and the appearance of unusually tight clustering in a whole cohort’s scores.
The fourth is the least visible and the most informative. A cohort whose scores bunch into a narrow band for no apparent reason is either genuinely alike or has been levelled by something in the administration, and that question is asked before the results are adopted rather than after.
What this page does not establish
We did not find, in our sources, coefficients setting a level at which a selection instrument’s stability becomes acceptable, nor figures attributed to any named instrument. Figures circulating outside our sources are not reproduced here.
The page also does not establish which of the four forms is reliability in the general case, nor a ranking of strength among them. What is set out above is that the form is chosen according to the shape of the use, and anything beyond that needs a source to settle it.
The figures in the worked example are assumed in order to show the structure. They are not the results of any existing instrument and are not attributed to one.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.