Qoyod National Day offer: up to 50% off plans and add-ons · until 30 September See the details
Qoyod
Pricing
Qoyod
Pricing

Work Sample Test

Term in Qoyod's Business Glossary. Practical definition with examples from the Saudi market.

What a work sample test is

A work sample test is a short task taken from the job itself, given to candidates under identical conditions, whose output is judged against a written standard prepared before anybody has been seen.

Its distinguishing property is that the candidate performs rather than describes. What it produces is an actual piece of work: a reconciled account, a reply to a customer complaint, a schedule built from raw figures.

Why the content is the argument

This instrument defends itself with its own content. If the task is drawn from the job in roughly the proportion it occupies there, the connection between what was measured and what will be done is visible before any calculation is performed. That is why it suits an organisation hiring in small numbers: its argument does not need cohorts or a year of waiting.

The same fact imposes a limit. A task that is not performed in the job destroys the argument entirely. A hard programming exercise lifted from a competition and bearing no resemblance to the role measures something else and keeps the name.

There is a finer point here that gets lost, and it is where the instrument is most often overread. The argument covers the task that was set, not the job as a whole. An organisation that set one account reconciliation has evidence about reconciliation and no evidence at all about the following up, the correspondence and the monthly close that fill out the rest of the role. Calling the result “suitable for the job” extends the claim past what was measured. Calling it “performed this task at this level” is what the evidence will carry.

How one is built

  • The task is drawn from the job’s frequent and consequential work, not from its rare work and not from its hardest.
  • The time is cut to what is needed to observe the performance. A candidate is not asked for what an employee produces in a full day.
  • Conditions are held constant: same time, same tools, same information given to everybody.
  • The marking key is written before administration and describes what counts as weak, acceptable and strong output.
  • Two independent assessors mark where that is possible, so that differences in how the key is read surface in the key itself.

There is a sixth item that is regularly skipped: deciding, before administration, what will be done with the score. Making it a filtering stage that drops everybody below a line is a different instrument from making it material read alongside everything else in the final decision. Both arrangements are legitimate. Confusing them is how a score gets used as a filter when it was gathered as discussion material.

A marking key is a set of weights, on a worked example

A key is not a list of desirable qualities. It is a distribution of weight across the things being looked at in the output. The example below is assumed in order to show the structure, and its numbers are design choices rather than a standard to copy.

Let the task be reconciling a bank account from a statement and a set of entries, and let the key carry three criteria whose weights total 100: the final reconciliation being correct, weighted 50; finding the open item planted in the data, weighted 30; and the output being clear enough to review, weighted 20. Each criterion is rated from 0 to 4, and its score is the rating multiplied by its weight and divided by 4.

A candidate rated 4, 3 and 2 by the first assessor scores 50, then 22.5, then 10, for a total of 82.5. The second assessor rates the same paper 4, 1 and 2, giving 50, then 7.5, then 10, for a total of 67.5.

The two readings are 15 points apart, and the whole of that 15 sits in one criterion. The other two criteria agree exactly. This is precisely what is lost when only the total is kept: two distant numbers look like a broad disagreement, when in fact they are a disagreement about the meaning of a single line in the key, which is repaired by describing that line more precisely and is not repaired by averaging the two totals.

So the rule is to store the rating for each criterion separately and never the total alone. The total tells you there is a disagreement. Only the breakdown tells you where it is, and that is the same discipline test reliability depends on for anything marked by a human.

The length of the task is decided by the cost of reading it

What kills this instrument in smaller organisations is not its design but its cost after design, and that cost is calculated before the task is chosen rather than after.

Suppose 30 candidates sit a task lasting 45 minutes, and reading and rating one paper takes 15 minutes, and two assessors mark. The total is 30 times 15 times 2, which is 900 minutes, or 15 hours of assessor time.

Now make the task three hours instead of 45 minutes. Reading one paper rises to roughly 40 minutes, and the total becomes 30 times 40 times 2, which is 2,400 minutes, or 40 hours.

The difference between those two figures decides what actually happens. A team without 40 hours will cut something, and the cut almost always falls on one of two places: the second assessor is dropped, or the key collapses into a general impression. Each of those removes exactly the element that made the instrument work.

So a short task is not a concession on rigour. It is the condition under which the key and the second assessor survive. An organisation that chooses a long task and then drops both has moved to a different instrument and kept the name. The assessor hours involved are internal recruiting time and belong in cost per hire like any other internal hour.

How it differs from a structured interview

A structured interview fixes the questions, their order and the rating scale, and its material is conversation. The material here is delivered output.

An interview tells you how a candidate thinks about what they are asked. A work sample shows you what they produce when left to it. The difference is not one of rigour, since both are governed by a written standard. It is a difference in the kind of material being judged.

The question that decides whether you ran one

Was the task set up for measurement, with uniform conditions and a key written in advance? Anything that fails that question is not a work sample, however strong an impression it left. What is worth knowing is that the arrangements sharing the name fail it in different places, and one of them does not fail it at all.

A trial working day fails on the first half. The candidate comes to the premises and does live work alongside the staff, which is output the organisation benefits from in its ordinary course rather than a task prepared in order to be measured, and calling it a test conceals that work was taken. The probation period fails on placement rather than on design: its place is after contracting rather than before, and its material is performance in the role itself over a continuous period. It is a regulated matter under the Saudi Labor Law (نظام العمل) with its own rules, set out in the probation period in the Saudi Labor Law, and this page does not deal with it.

A job knowledge test passes the question and is still not this, which is the case worth holding on to. It is prepared in advance, its conditions are uniform and it is marked against a key. What it collects is an account in words of what the candidate knows, and the whole argument of this instrument is that the account and the work come apart: somebody who can say how a reconciliation is built may not build one, and somebody who built one has already shown you the building. So the question above rules things out reliably and does not on its own rule anything in.

What spoils it

  • A long unpaid task, which excludes whoever does not have free time and reintroduces bias through availability rather than ability.
  • Real work disguised as a test, where a job the organisation genuinely needs is completed free under the heading of assessment.
  • Information given to one candidate and not another, which turns the comparison into one between two briefings rather than two performances.
  • Marking after the author is known, in cases where concealment was possible.
  • Reusing one task for years until it circulates, at which point it measures who saw it first.
  • Marking each paper against the one before it rather than against the key, which lets the ruler drift with the reading order and penalises whoever follows a strong paper.

The last of those has a simple procedural fix: read criterion by criterion rather than paper by paper, rating the reconciliation criterion across every paper before moving to the next. That holds one thing in mind at a time while judging.

Edge cases, and who they bind

The rules above are written for the ordinary case. Four situations change the reading, and each is about a particular role rather than about the instrument in general. In entry level roles the candidate has not performed the work yet, so what is measured is closer to speed of learning than to mastery, and naming the score accurately prevents a judgement being made in the wrong place. Where the material is confidential, as with real customer data, the task is built on manufactured data resembling it in shape, size and difficulty rather than in content. In remote work the task is performed unsupervised and there is no way to know who performed it, so it is used as a first stage and its author is then asked about their choices face to face, which separates whoever did it from whoever copied it. And where a candidate needs a different arrangement in order to perform, it is the conditions that are adjusted while the thing measured and its key stay as they were, because adjusting the key itself produces two scores that cannot be compared.

What the test produces is personal data

What a candidate submits, and what is written about them in the key, relate to an identified person, so the Personal Data Protection Law (نظام حماية البيانات الشخصية), issued by Royal Decree M/19 of 9/2/1443H and amended by Royal Decree M/148 of 5/9/1444H, applies to them.

Several of its duties bear directly on this. Article 11(1) of the Personal Data Protection Law requires the purpose of collection to relate directly to the controller’s own purposes and not to conflict with any provision established in law. Article 10 of the same law permits processing only for the purpose the data was collected for, subject to the exceptions it lists. Article 11(3) of that same law requires the content of the data to be adequate and confined to the minimum necessary for that purpose. Article 19 of the Personal Data Protection Law requires the controller to take organisational, administrative and technical measures to safeguard the data, including when it is transferred. Article 18 of that same law requires destruction once the need has ended, with retention only on a legal basis or for a pending case. Those are among the law’s duties rather than all of them.

The practical effect is specific. A test paper was collected for a particular hiring decision, so once that decision has passed it is not carried across to another purpose, such as becoming internal training material or a worked example shown to somebody other than its author. A new purpose is governed on its own terms.

Two further practices are recommended in this area as practice rather than drawn from the text of the law: telling the candidate before administration what they will be asked to submit and who will see it, and keeping access confined to those taking part in the decision.

Its ceiling

What can be observed in an hour is what this instrument is good for. Anything that shows itself only over a long acquaintance, such as whether quality holds up through a pressured season, is beyond a single task however well built.

The same goes for anything whose material is time: building a relationship with a difficult customer, or carrying a long piece of work through to the end as its priorities shift. These do not compress into a task, and presenting them as one produces a measurement of something easier that has been substituted for them.

What this page does not establish

We did not find, in our sources, a provision setting out a rule on an employer running selection tests in the Saudi private sector: not on the conditions for administering one, not on its limits, not on payment for a candidate’s time, and not on how long its papers are kept or how they are destroyed. This page therefore sets out none of that, and what is described above under practice is labelled as practice. That is a statement about our sourcing and not about the statute book.

The page also does not establish how far this instrument predicts anything, nor a ranking for it among selection instruments, nor a score at which it becomes acceptable. Figures circulating outside our sources are not reproduced here, and the cut score in any given application is a choice the organisation writes down, with its reason, before it sees the scores.

Qoyod HR

A standalone Saudi HR system

One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.

Explore Qoyod HR

A standalone system on its own subscription. The connection to Qoyod Accounting is now available.

Related terms

Ready to apply accounting the right way?

Qoyod runs your accounting with precision and full ZATCA compliance

Try Qoyod free for 14 days — No credit card required.