What test validity is
Test validity is the extent to which evidence supports a particular inference drawn from a test score for a particular purpose. It answers one question: is the thing this score measures the thing we want to decide with?
The wording matters more than it looks. Validity is not a property the test carries around with it, and it is not a certificate an instrument arrives holding. It is a claim about a score, a decision, and the relationship between them, which means it can be true in one organisation and false in the one next door using the same paper.
Why no test is valid in general
A test that measures spreadsheet proficiency may well support a decision to hire a junior accountant. The same test, with the same questions and the same key, does not support a decision to promote somebody into leading a team. Nothing about the paper changed between those two uses. What changed is what the score was asked to carry.
So the only usable form of the question is: valid for which decision? An instrument described as valid without a decision attached has not been described at all.
This has an immediate practical consequence that organisations discover late. Moving a working test from one role to another moves the questions and leaves the evidence behind. Whatever was established in the first role has to be established again in the second, because the inference is new even though the instrument is not.
Write the inference as a sentence before you look for evidence
The single most useful thing to do in this area costs nothing: write the inference out as a full sentence, before anybody discusses evidence at all. A written sentence exposes in one line what the phrase “it is a good test” is hiding.
The sentence names five things: the score, the threshold, the person, the decision, and the period the decision covers if it covers one. For example: a candidate scoring 70 out of 100 on the reconciliations test can prepare a monthly bank reconciliation without review by a senior accountant. A reader can interrogate every clause of that. Where did 70 come from? What in the test corresponds to preparing a reconciliation? How do we know that somebody who reaches it stops needing review?
Now set it beside the sentence most organisations actually use: we use this test to screen applicants. There is no threshold in it, no specific decision, and nothing in it that could turn out to be wrong. An inference that cannot be wrong cannot be evidenced either, and the two facts are the same fact.
Writing the sentence also makes the threshold visible as a claim rather than an administrative setting. Raising a cut score from 70 to 85 does not make the existing inference stronger. It makes it a different inference, with its own evidence to find. An organisation that lifts its threshold because applicant volume rose has changed what it is asserting while believing it has tightened a process.
Decision first, then inference, then evidence route
That order is reversed in practice far more often than it is followed. The usual sequence is that a tool is bought or borrowed, and a use is then found for it. The defensible sequence runs the other way:
- Name the decision. What will actually be done with the score? Initial screening, ranking a shortlist, a direct offer, or routing somebody into a training track.
- Write the inference that would justify doing that, in the form above.
- Choose the evidence route that suits what is being measured and the data the organisation can realistically get.
Running it in this order exposes a tool with no place before any money is spent on it. An organisation that cannot name the decision needing the score does not need the instrument, and one that names the decision but cannot write the inference has a problem in how it has understood the decision rather than a problem with the tool.
Three routes to evidence, not three kinds of test
The literature talks about “types” of validity, and the word does real damage. They are not categories a test is sorted into. They are different ways of gathering evidence about the same inference, and a single instrument can legitimately be supported by more than one at once.
- Content evidence: a judgement about what is in the test, compared against what is in the job.
- Predictive evidence: a comparison between the score and what later showed up in the work.
- Construct evidence: how the score itself behaves against other measures.
Which route fits is decided by the instrument and its purpose rather than by preference. A work sample test that puts a piece of the job in front of the candidate argues from its content first. An instrument claiming to measure a trait nobody can observe directly, which is the position integrity testing is in, gets nothing at all from a content argument, because there is no visible job content to compare it against.
Treating the three as categories has a practical cost, not a verbal one. Whoever reads them as types assumes the instrument has to be one of them, picks a label, and stops. Whoever reads them as routes understands that two routes can support one inference, and that a weak route is not repaired by a confident label.
The evidence is about your inference, and it does not arrive with the product
This follows from everything above and it saves organisations real money. What ships with an instrument is documentation about the instrument. It is not evidence for your inference, because your inference is about your role, your threshold and your decision, and the vendor has never seen any of the three.
The burden therefore stays with whoever uses the score and does not transfer to whoever produced it. An organisation using an instrument for a decision it was never studied for owns that outcome, and a technically sound manual sitting in the drawer does not change it.
There is a single question test to apply before purchase. Does the documentation contain an inference resembling mine, in a role resembling mine, at a threshold resembling mine? If the answer is no on any one of the three, what you are holding is material to reason from, not support to rely on.
What it needs from reliability, and what reliability cannot give it
Test reliability is necessary here and nowhere near sufficient. A score that moves for no reason will not carry an inference at all, because there is nothing stable enough to build one on. A perfectly stable score can be perfectly stable while measuring something with no bearing on the job.
The order between them is an order of inspection: look at whether the score holds still, and if it does, look at what it means. Stopping after the first step is the commonest error in this whole area, for the least interesting reason. Stability can be examined with numbers the organisation already holds, and meaning needs either a judgement or a wait.
What invalidates an inference that was previously supported
Validity is lost, not acquired once. Six events end it, and they are listed in rough order of how visible they are:
- The job changed and the test did not. The evidence was gathered on work that is now different work.
- Administration conditions changed: more time, a tool that was not previously allowed, or a move from a supervised room to a remote sitting.
- The questions circulated, so the score became a measure of prior exposure.
- The score was used for a decision it was never studied for.
- The threshold moved without support, upward when applicants were plentiful or downward when they were scarce.
- The applicant population changed markedly, as when an advertisement opens onto a labour market it had not reached before.
The fourth is the dangerous one precisely because it leaves no trace. The instrument did not change, the questions did not leak, the job is as it was, and the score simply migrated from one decision to another in the course of a single meeting. The cheap guard is to write the decision the instrument was prepared for onto the instrument’s own report, so that anyone reaching for it for something else reads the limit before they use it.
When several instruments feed one decision
Most selection processes do not rest on one score. They run two or three instruments in sequence, and that raises a question which is almost never asked out loud: what is the inference about? Each instrument separately, or the combined result?
The answer follows from how the instruments were arranged. If they are sequential filters, with candidates dropped at each threshold, then each instrument carries its own inference and answers for itself, and whoever fails the first never reaches the second. If the scores are combined into a single total, there is one inference about the total, and the weight given to each instrument becomes part of the claim rather than an administrative detail.
That is where weighting shows its teeth. An organisation giving a structured interview 60 percent of the total and the practical task 40 percent may be entirely right, but it chose that split, and asking what supports it is a fair question that is not answered by pointing out that the two add to 100. The common pattern in practice is that weights are set to bring the total close to a conclusion already reached, which turns the combined score into a numbered picture of an opinion that predated every instrument in the process.
What gets offered as evidence and is not
Several things are regularly presented in support of validity and support something else instead. Candidate acceptance, meaning that applicants found the instrument reasonable and fair, is worth having for its own sake and is part of candidate experience, but it says nothing about whether the score supports the decision. Cost, reputation and the number of organisations using an instrument are evidence of adoption. Polished output, with detailed reports and comparison charts, describes the presentation layer rather than the support underneath it.
What they share is that each can be checked in an afternoon. Evidence for an inference cannot: it has to be assembled against a role, a threshold and a decision that are yours alone, and on the predictive route it cannot begin at all until the outcome has arrived. That asymmetry in what it costs to look, rather than any confusion about what the words mean, is what keeps putting these in the position where evidence belongs.
What this page does not establish
We did not find, in our sources, figures establishing how valid any particular selection instrument is, nor a ranking among instruments, nor a threshold at which a test becomes acceptable. Figures circulating outside our sources are not reproduced here.
This page also does not set out any regulatory position on an employer running selection tests in the Saudi private sector. Our sources record a search for such a rule that did not find one, which is a statement about our sourcing rather than about the statute book. Anyone needing an answer there should go to the competent authority and its own published document.
The thresholds and figures used in the example inference above illustrate the shape of the sentence. They are not recommended values and are not taken from a source.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.