What integrity testing is
Integrity testing covers instruments administered to candidates in order to estimate the likelihood of harmful conduct at work: taking what is not theirs, breaching a published rule, or reporting something other than what occurred when asked.
Its distinguishing property is that it does not measure a capacity to do a job. It measures an estimate of the likelihood of conduct that has not yet occurred, and that fact alone makes its output read differently from the output of every other selection instrument. Its subject is one named attribute, which is what separates it from a situational judgement test, whose subject is conduct in a composite situation that the attribute may or may not enter.
Start with the base rate, because it decides how the output reads
The thing most often overlooked when these instruments are read is not their accuracy. It is the rarity of what they are looking for. The calculation below shows the structure; every figure in it is assumed for that purpose and is not attributed to any existing instrument or to any measured incidence.
Assume 1,000 candidates. Assume that 50 of them would engage in the harmful conduct if hired, which is 5 percent, and therefore that 950 would not; those two groups account for all 1,000 and nobody is in both. Assume an instrument that flags 80 percent of those who would, and wrongly flags 10 percent of those who would not.
- Of the 50 who would engage in it, the instrument flags 40.
- Of the 950 who would not, the instrument flags 95.
- So 135 candidates are flagged in total, and that 135 has exactly those two groups inside it: 40 who would have engaged in the conduct and 95 who would not.
Of the 135 candidates the instrument flagged, 95 would not have done anything. That is 70.4 percent of the flagged, and it is not a defect in the assumed figures. It follows directly from what is being looked for being rare: a small error rate applied to a large group produces more people than a high hit rate applied to a small one.
The rarer the conduct, the harder this bites. Hold the instrument’s two rates and move the base rate to 1 percent, so 10 of the 1,000 would engage and 990 would not. The instrument now flags 8 of the 10 and 99 of the 990, giving 107 flagged candidates, of whom 8 match what it was looking for. Of those 107, 99 would not have done anything, which is 92.5 percent of the flagged.
The rule that comes out of this is plain: the output of these instruments is not read as a detection, and an exclusion is not built on a flag alone.
Where the cut sits, and who chooses it
Connected to the arithmetic above is the choice of the cut. The instrument produces a continuous score, and the decision built on it is binary: the candidate proceeds or the candidate is dropped. The point at which the first becomes the second is not a property of the instrument. It is a choice the organisation makes, and raising it by a single point changes how many people are dropped without changing anything at all about what was measured. The same distinction is worked through numerically on the cognitive ability test page, where a ten point move in the line is costed against a cohort.
So the cut is written down, with its reason, before administration rather than after the distribution has been seen. Choosing it after the scores are in front of you makes the rule follow the result somebody wanted.
Two forms it is written in
- Overt. It asks the candidate directly about their attitudes and opinions on what counts as a breach, and sometimes about what they have done in the past.
- Indirect. It asks about general traits and preferences with no mention of any breach, and the answers are then read through a key that links them to the estimated conduct.
Both forms arrive at a score read under one name, and the routes by which they arrive are different. An organisation that bought an instrument without knowing which form it uses does not know what its candidates were asked.
Each form has a weakness peculiar to it. With the overt form the candidate can see what is wanted, so answering becomes a selection rather than a disclosure. With the indirect form it is the organisation that cannot see what was linked to what, so it carries a score it cannot explain to anybody who asks about it, and cannot review if it comes to doubt it.
Why the burden of proof falls on construct evidence
Integrity is not a task that can be set and performed. It is a concept posited behind conduct, so nothing is gained here by the argument from matching content to the job, which is the argument a work sample test relies on.
The question put to whoever supplies the instrument is therefore the one test validity frames as construct evidence: by what examinable thing was it established that this score connects to the concept printed on its cover? A descriptive brochure does not stand in for that.
The faking problem, and why the two remedies fall short
The instrument is administered in a situation where the candidate has an obvious interest. They know which answers are wanted, and they need no special insight to choose them.
That is where the caution sits. A high score may mean uprightness, and it may mean a good reading of the situation. The instrument does not separate the two by itself, and one consequence of that is that the candidate whose score comes out lowest is the one who read the situation least well.
Two remedies are tried and neither reaches the goal. The first is warning the candidate that the instrument contains means of detecting a manufactured answer, which moderates the exaggeration among those who believed the warning and therefore widens the gap between them and those who did not. The second, which is the more useful, forces a choice between two statements both of which are desirable or both of which are not, so that no safe obvious answer is available. Even then, a candidate who knows what the instrument measures can still weigh the options, and the comparison has become one between two choices rather than between two levels.
The practical conclusion is that faking is reduced and not removed, and that every reduction costs something in the clarity of the instrument or in the time it takes. Whoever reads a score of this kind is reading a number containing both the effect of conduct and the effect of reading the situation, with nothing available to separate them.
How it differs from a background check
A background check verifies what the candidate stated about themselves: a qualification that was or was not issued, a period of employment that did or did not happen. That deals with past events capable of proof; this deals with an estimate of what has not occurred.
The procedural consequence is the one that matters at the point of decision. Somebody who disputes what a background check attributes to them can produce a document. Somebody whose score came out low here does not know what is being attributed to them at all, so there is nothing for them to answer. The first has a route to review, and the second has nothing but a retake, which will produce another score close to the first.
What may not be borrowed from occupational fitness
An arrangement from a different subject is sometimes cited in support of these instruments: the occupational fitness examination governed by the regulation on occupational fitness examinations and noncommunicable diseases (لائحة فحوصات اللياقة المهنية والأمراض غير المعدية), issued by Ministerial Decision 33232 of 11/3/1447H. The borrowing does not hold, and the reason it does not is finer than the two subjects simply being different.
That regulation erects, in its Article 17(4), a wall between whoever conducts the examination and the employer: only one of three determinations reaches the employer, fit, fit with restrictions or considerations, or unfit. The point of the wall is to keep clinical detail away from somebody who has no entitlement to see it.
Carrying that arrangement across to an integrity instrument the employer runs itself inverts the protection into a route. It takes a provision written to stop the employer reaching the detail and uses it to license an instrument the employer administers and whose output it reads in full. The two examinations differ in their subject, in who conducts them and in who sees the result, so the rules of one are not carried to the other in any respect.
A score is not an event
What is produced here is an estimated likelihood, not proof of an act. Building an exclusion on a score alone treats an estimate as though it were an event.
What actually happens after hiring has its own territory: internal investigation and disciplinary procedure, both of which rest on a specified event and a known process. The logic of these instruments is not carried into that territory, and the logic of that territory is not carried back here.
The most damaging form of that transfer is keeping the score in the employee file after the person is hired, where it is read two years later at the first suspicion and used to tilt a judgement about an event the instrument knew nothing of and was never built for. The score is an estimate made before hiring, and the passage of time does not turn it into evidence of anything.
The candidate’s data in this area
What is collected here is personal data to which the Personal Data Protection Law (نظام حماية البيانات الشخصية), issued by Royal Decree M/19 of 9/2/1443H and amended by Royal Decree M/148 of 5/9/1444H, applies. Among its duties are specifying the purpose of collection under Article 11(1) of that law, confining the content of the data to the minimum necessary for that purpose under Article 11(3) of the same law, and taking organisational, administrative and technical measures to safeguard what is collected under Article 19 of that same law.
Alongside those, and recommended as practice rather than drawn from the text of the law: telling the candidate about the instrument before it is administered, keeping access to its result confined to whoever takes the decision, deciding from the outset where the result is held and who may open it, not circulating a score in general messages inside the organisation, and knowing what becomes of it once the application has ended, whether or not its subject was hired.
Design controls do more than screening at the gate
More useful than any of this is what comes before it: controls in the design of the work that make harm difficult for anybody who wanted to cause it. Divided authorities, a separation between whoever disburses and whoever reviews, and a published code of conduct stating what counts as a breach before one occurs.
The advantage of that route is that it assumes no knowledge of individuals. A control preventing disbursement without review operates on somebody who intended a breach and on somebody who made a mistake, and it keeps operating after whoever holds the post has changed. Screening at the gate stops having any effect on the day of appointment.
What this page does not establish
We did not find, in our sources, any particular instrument or any body issuing one, nor figures establishing how far these instruments predict anything. Figures circulating outside our sources are not reproduced here.
Nor did we find a provision setting out a rule on the use of these instruments in the Saudi private sector, so this page sets out nothing on that. Anyone needing an answer there should go to the competent authority and its own published document.
The page also does not establish a period for which the result of one of these instruments is kept, nor a procedure for destroying it, nor who may see it as a matter of text. What is described above in that area is practice, and anything past it needs a provision to settle it that we did not find.
All the figures in the calculation above are assumed, in order to show the effect of rarity on how a flag reads. They are not the properties of any instrument and are not a measured incidence in any labour market.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.