Qoyod National Day offer: up to 50% off plans and add-ons · until 30 September See the details
Qoyod
Pricing
Qoyod
Pricing

Situational Judgement Test

Term in Qoyod's Business Glossary. Practical definition with examples from the Saudi market.

What a situational judgement test is

A situational judgement test presents the candidate with a situation from the work, sets out several possible courses of action alongside it, and asks them to choose among those courses according to a written instruction. Their answer is then compared against a key prepared before administration.

Its distinguishing property is that what is measured is a choice among options placed in front of the candidate. It is not output they hand over and it is not an account of something that happened. It asks what would be done in a situation that has not occurred, and the options are supplied in full, so the candidate is never asked to produce a course of action of their own.

It is sometimes called a low fidelity simulation, meaning that it presents the situation by description rather than by performance. That label is useful because it says in one phrase where the instrument sits among the others: cheaper than anything performed, and further than anything performed from the work itself.

The key is the instrument, and the key is what gets made

The common assumption is that the work here is writing the situations. Writing the situations is the easier half. The half the instrument actually stands on is how the correctness of a choice is settled, and there are two routes to that in practice and no third.

  • A key from subject matter experts. The situation and its options are put to a number of people who do the work and do it well, each rates how effective every option is, and the key is built from the average of their ratings.
  • A key from outcomes. The instrument is administered to people currently in the role, and the options chosen by those whose performance is regarded as good are identified. This needs numbers and time, and is not available to an organisation hiring in ones and twos.

The first route is the common one, and it contains the failure that occurs most often: taking the average of the expert ratings and discarding their spread. Take a situation with four options, rated by five experts on a scale from 1 to 5.

  • First option: ratings of 5, 4, 5, 5 and 4. Mean 4.6, and the ratings span one point.
  • Second option: ratings of 3, 3, 4, 2 and 3. Mean 3.0, and the ratings span two points.
  • Third option: ratings of 2, 2, 1, 2 and 2. Mean 1.8, and the ratings span one point.
  • Fourth option: ratings of 4, 5, 3, 4 and 5. Mean 4.2, and the ratings span two points.

Read as means alone, those four rank cleanly. The distance between the first option and the fourth is 0.4 of a point on a five point scale, and the fourth option’s ratings are spread across twice the range of the first’s. What that says is that the experts themselves did not agree about the fourth option, while they did agree about the first. The key nevertheless records the first as right and the fourth as wrong.

Hence a working rule: an option on which the experts span two points or more does not carry a difference in the score. Either it is rewritten until they agree, or the whole situation is dropped. An option that people who know the work disagreed about is not one a candidate who has never done the work can reasonably be marked down on.

What a usable situation is made of

A situation that produces a difference between candidates has three properties, and losing any one of them makes the question answerable without thinking.

  • A conflict between two things both of which are wanted. Something due on a deadline, and on the way to it a matter that touches a customer. A situation with one obvious right answer and three obvious wrong ones separates people who read the language from people who know the work, and nothing else.
  • A statement of who holds the decision. Some options suit a person with authority to act and do not suit a person without it. A situation that does not say where the candidate stands in the chain of authority gets answered according to whatever each respondent assumed.
  • Sufficiency of what is given. A situation needing information that was not supplied measures what the candidate assumed rather than what they chose, and the key changes with the assumption.

The source of these situations is the work rather than the imagination: they are collected from events that recurred in the role and reviewed with the people who perform it. That is the same route by which behavioral interview questions and a work sample test are built, even though what each extracts from the material is different.

Prior exposure does more here than anywhere else

This is the selection instrument most affected by familiarity with it, for two reasons. Its material is text that can be read and remembered, and what it asks for is knowledge of what counts as the right answer rather than a capacity built over years. A candidate who has sat two of these enters the third already knowing the structure of the question and what the key is likely to reward.

The consequence for reading a score is that a high one is consistent with two different things: knowledge of the work, or knowledge of the instrument. Nothing inside the instrument separates them, and only evidence of a different kind can. That by itself is sufficient reason never to settle an appointment on this instrument alone. The consequence for fairness is that whoever has been through many hiring processes arrives with an advantage unrelated to the role, and it is an advantage that shows up in no report.

The instruction wording changes what is measured

The line written above the situation is not a detail of phrasing. It decides the construct.

What should be done? measures knowledge of what counts as right in this work. It is learnable, and whoever read the organisation’s policy before sitting the test scores higher on it.

What would you do? measures what is likely actually to happen. It is closer to what is wanted and more exposed to the candidate answering with whatever they believe is preferred.

The difference is not theoretical. A candidate who knows the right answer and does not act on it produces two different scores under the two wordings, and that candidate is precisely the one the process is trying to identify. Mixing the two wordings within one instrument collects two scores measuring two things and then reads them as one number.

The scoring method, and what it throws away

The common arrangement asks the candidate to pick the most effective option and the least effective one. That narrows the effect of guessing: a candidate choosing at random from four options hits the most effective one alone 25 percent of the time, and hits both the most and the least effective together in one case out of twelve, which is about 8.3 percent.

Its price is that it makes the key binary. The candidate who picked the fourth option in the example above is recorded as wrong, and the same experts rated that option 4.2 out of 5. The score says they erred; the key itself says their choice was close to the most effective one.

The alternative is to ask for a rating of every option on a scale, and to score the candidate by their closeness to the experts’ ratings rather than by matching a single choice. That takes longer to administer and produces a finer output, and choosing between the two is a design decision the organisation takes and writes a reason for. It is not a property of the instrument.

How it parts from the instruments next to it

A work sample test produces a piece of work the candidate makes, and its distinguishing property is that the candidate performs rather than describes. This instrument keeps them describing: they read and they choose. The first shows you what they produce when left to it; this shows you what they favour when the alternatives are laid out for them.

Behavioral interview questions take their material from situations that actually occurred in the candidate’s past, and this instrument’s material is a situation that never occurred. So the first does not work for somebody who has not yet done the work, and this one does, because it asks for no record.

A cognitive ability test measures reasoning from material inside the question, needing no background in any particular field. This rests on knowledge of a work context: the most effective option in an administrative situation is not the most effective one in a field situation, and the key is drawn from that context.

And integrity testing has one named attribute as its subject, where this has conduct in a composite situation that attribute may or may not enter.

What spoils it

  • Situations lifted from a general source. The key then comes from somebody else’s work, and what counts as most effective in one organisation may be the worst option in another, because authority is distributed differently.
  • An obviously poor option. A choice nobody would take shrinks the question to three options and raises the guessing rate, while leaving the length of the test exactly as it was.
  • Reusing the same bank for years. The situations circulate, and the instrument comes to measure who saw them first. They circulate faster than other material because they are memorised by reading.
  • Language heavier than the language of the work. A long situation in complicated phrasing measures reading speed and keeps the name.
  • Making it an absolute barrier. The point here is the same one made about every instrument of this kind: the instrument ranks, and the line excludes. What moving that line costs is worked through numerically on the cognitive ability test page.

Where it earns its cost, and where it does not

Its usual place is after the first screen and before interviewing, where a large number is to be narrowed at a cost per candidate that barely moves with volume. It is used most in roles where applicants resemble one another in qualification and differ in how they act, such as customer facing and field service roles.

Its cost is not in administration but in building the key. The experts’ time spent rating options is spent once and then spread across everybody tested afterwards. An organisation testing five people a year spends the whole of that time on five. An organisation testing two hundred spreads it across two hundred. So the question is settled by how many will be tested rather than by how good the instrument is, and the usefulness of any instrument is measured partly by the room in front of it rather than by its precision alone.

What this page does not establish

We did not find, in our sources, anything establishing how far these instruments predict performance, nor their stability, nor a numerical ranking among selection instruments, nor a recommended number of situations or options, nor a score at which one becomes acceptable. The figures in the example above are assumed in order to show how the calculation runs.

This page does not endorse any particular instrument or any body issuing one, and it does not establish whether anything in Saudi Arabia regulates the use of selection tests in the private sector. Our sources record a search for such a rule that did not find one, which is a statement about our sourcing and not about the statute book.

The scores and answers this instrument produces are personal data. The duties that follow from that are set out on the cognitive ability test page and are not repeated here, and this page states none of them.

Qoyod HR

A standalone Saudi HR system

One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.

Explore Qoyod HR

A standalone system on its own subscription. The connection to Qoyod Accounting is now available.

Related terms

Ready to apply accounting the right way?

Qoyod runs your accounting with precision and full ZATCA compliance

Try Qoyod free for 14 days — No credit card required.