Qoyod
Pricing
Qoyod
Pricing

Predictive Validity

Term in Qoyod's Business Glossary. Practical definition with examples from the Saudi market.

What predictive validity is

Predictive validity (الصدق التنبؤي) is the extent to which a score taken before hiring agrees with an outcome that appeared in the work afterwards. It is evidence gathered over time: something is measured today, the organisation waits, and the measurement is then compared with what happened.

It is one of the three routes to evidence set out under test validity. What sets it apart from the other two is that it alone cannot be settled until some time has passed after the hire.

The criterion is half of predictive validity

The second side of the comparison is called the criterion (المحكّ): the outcome against which the prediction is checked. Choosing it is a decision that comes before any calculation, because the whole of the evidence changes when the criterion changes.

A tool can agree well with a performance rating in the first year and not agree with staying in the job, and the reverse can also happen. Anyone who says “this test predicts” without saying what it predicts has said nothing that can be checked.

Some of the criteria in use are those set out under quality of hire. As an indicator, quality of hire measures the outcome of the hires themselves; in a predictive validity check it can serve as the second side of the comparison, not as a verdict on the tool.

Two faults in the criterion that distort predictive validity

A criterion is a measurement like any other, and it is open to the same faults. Its two opposite faults deserve separate names, because they push the interpretation in opposite directions:

  • A deficient criterion. It covers only part of the outcome that matters. A performance rating built on sales volume alone leaves out the quality of the customer relationship and the accuracy of the records, and the tool may predict what was left out rather than what was kept.
  • A contaminated criterion. It takes in things that are not the employee’s performance. A rating swayed by a rich sales territory, a strong team or a lenient manager measures the circumstances along with the person, so the tool is blamed for failing to predict something that was never performance in the first place.

A criterion based on staying in the job has a peculiarity of its own. For anyone still employed, it has not happened yet. Someone who has completed a year without leaving is not someone who stayed; they are someone who has not left so far. A check built on this kind of criterion has to fix the date of observation for each case and compare intakes over equal periods. Otherwise it compares people who have had the chance to leave with people who have not yet had it.

The first remedy for both faults is to write the criterion down before its data is collected, in the same terms by which everyone in the study will be measured. A criterion defined after the results have been seen ends up described in a way that fits them.

How a concurrent check differs from predictive validity

A closely similar arrangement can be confused with predictive validity: the test is given to current employees today, and their scores are compared with their current performance ratings. It is quicker, since there is no wait.

What it gathers, however, is not prediction. The score and the criterion were taken at the same moment from people who have already spent time in the work, so any agreement may be the effect of experience on the score rather than the effect of the score on performance. An employee with three years in a role answers questions about the role as someone who has done it, and that is not the position a candidate is in.

This does not make the concurrent arrangement useless. It helps in examining the instrument itself and in finding items on which everyone scores alike, but it is not offered as evidence that the score comes before the outcome and points to it.

Why predictive validity is measured only on the people hired

An organisation sees in the work only the people it accepted, and if the score fed the hiring decision, those people come from the upper part of the score range. The comparison therefore runs over a narrow slice of the range, not over every applicant. The selection ratio describes how narrow that slice is.

The effect is that agreement appears weaker than it is, because the large differences were removed before anything was measured. This is a structural property of any study of this kind in which the score fed the decision, not an error that collecting more data on the people hired would put right.

A worked example of range restriction in predictive validity

Forty people applied for a role, and their test scores ranged from 31 to 88, a range of 57 points. Twelve of them were hired, with scores between 62 and 88, a range of 26 points. The range over which the comparison runs is less than half the range over which the test was given.

The twelve were then ordered by their test scores, and ordered again by their performance ratings after a year. The two orderings matched in 8 positions out of 12. That looks like moderate agreement, and it can be taken as a weakness in the tool.

Consider what did not enter the calculation. The 28 applicants who were not hired, among them someone who scored 31 and someone who scored 44, were never given the chance to have their performance measured. Suppose the tool separates sharply between a candidate scoring 35 and one scoring 85: that very separation is what was removed from the data before anyone looked at it. The 8 out of 12 says something about the tool’s ability to discriminate within the accepted group alone, and nothing about its ability to separate the accepted from the rejected.

A rule of interpretation follows. Weak agreement calculated on the people hired alone is a lower bound, not a central estimate. Anyone who takes it as the amount the tool achieves has looked at half the data and judged the whole. The effect grows as the share accepted shrinks: an organisation that accepts one applicant in forty measures over a narrower slice than one that accepts one in three.

The figures in this example are illustrative. They show the effect of a narrow range; they are not the results of a study and are not drawn from any existing tool. In the sources we reviewed, we found no coefficient assigning a particular level of predictive validity to a named selection tool, no numerical ranking of tools by it, and no threshold at which agreement counts as sufficient. We do not repeat figures of that kind that circulate outside them.

The interval between measurement and criterion in predictive validity

Whatever happens during that interval mixes into the result. The new hire passes through a manager, employee onboarding and a role whose content may change, and each of these affects what will later be recorded as performance.

So weak agreement has two possible explanations to weigh before any judgement: either the tool does not predict, or what happened after the hire decided the outcome. Attributing the result to the tool alone settles one of two possibilities without examining the other.

The length of the interval is itself a decision with consequences. A measurement after three months falls in the middle of learning, so it captures how quickly someone picks things up more than the settled level of their performance. A measurement after three years moves so far from the point of testing that what lies between the two outweighs what came before. A period of about a year is one possible choice, not because it is an established standard but because it can fall after the employee has settled and before the role changes. We do not fix any particular interval as the right one, nor a minimum number of cases at which the check becomes sound.

What spoils a predictive validity check

  • The rater knows the candidate’s test score when giving the later rating, so the agreement becomes an effect of that knowledge rather than evidence for it.
  • Small numbers. An organisation that hires a handful of people a year will not have enough cases on which to base a conclusion.
  • Changing the tool or its cut score during the period, so that the comparison is between two scores that do not mean the same thing.
  • A moving criterion. The performance appraisal form, or what its ratings mean, changed between one intake and the next.
  • The weakest leaving before the measurement. People who left in the first months drop out of the data, and they may be part of the picture the check is about.
  • Choosing the criterion after seeing the results, by trying one criterion, finding no agreement, and replacing it with another until one agrees.

The first of these can go unnoticed, and it is simple to prevent: the selection scores are kept away from whoever gives the performance rating, and stored somewhere opened only when the check is run.

What a single organisation can do about predictive validity

A full check needs data that a single organisation may not hold. What is within reach can be simpler, and still more useful than doing nothing: keep what was measured at selection, and review it after a year against what has since appeared.

Four things, at the least, make that review possible: the score of every applicant, not only of those accepted; the date of the measurement; the version of the test that was given; and the decision that was taken, with the name of whoever took it. An organisation that keeps the scores of the accepted alone has thrown away half its data on the first day without realising it.

An organisation that uses a tool for years without returning to it with this question is paying for a measurement without knowing whether it has helped. A periodic review of the cases, few as they may be, is worth more than a composite figure that nobody examines.

If the agreement turns out weak, three things are worth examining before the tool is dropped. Is the criterion what the organisation actually wants? Does what happened after the hire explain the result? Did the narrow range hide the difference? If all three are ruled out, a fourth possibility remains: the tool does not suit this decision.

This is an explanation of the concept, not legal advice.

Qoyod HR

A standalone Saudi HR system

One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.

Explore Qoyod HR

A standalone system on its own subscription. The connection to Qoyod Accounting is now available.

Related terms

Ready to apply accounting the right way?

Qoyod runs your accounting with precision and full ZATCA compliance

Try Qoyod free for 14 days — No credit card required.