What construct validity is
Construct validity is the extent to which the evidence shows that a test measures the concept it claims to measure, where that concept is something that cannot be observed directly.
It is the third of the routes to evidence set out under test validity. Its material is the behaviour of the score itself: how the score moves against other measures, not what the content of the questions appears to be.
The construct behind construct validity
A construct is a concept made to explain what we observe, not a thing anyone can point to. Motivation is an example, as are the ability to work under pressure and a disposition towards risk.
Nobody sees motivation. What can be seen are traces from which it is inferred: taking the initiative on a task, or persisting after a setback. The construct is the name given to whatever is assumed to lie behind those traces.
That is why this route to evidence is needed at all. A task from the job can be compared with the job. Motivation has nothing to be compared with except other measures and the behaviour it is expected to produce.
Construct validity starts with a written definition of the construct
The first requirement is a written definition of the concept, stating what falls inside it and what falls outside. Without that definition there is nothing for the instrument to hit or to miss.
A useful definition states three things: what is assumed to lie behind the behaviour, in which situations it is expected to appear, and which neighbouring concept it is not to be confused with. Take the concept of persistence. A definition might read: a tendency to carry on with a task after the first setback, appearing in situations where stopping offers an easy way out, and meaning neither speed of completion nor satisfaction with the work.
The third element is the one that keeps a neighbouring trait out, and it is also the one a definition can leave out. A writer who defines a concept without saying what it is not leaves room for a neighbouring trait to be measured and given the concept’s name, and that can happen in any instrument whose definition stops at the first two elements.
A test’s name is not evidence of construct validity
Whoever builds a test names it, and the name binds nobody. An instrument titled with a concept remains a set of questions until evidence is produced that its score is connected to that concept.
This gap can open in selection instruments bought ready made. The name is taken from the brochure and a decision is built on it, while nobody has asked the supplier for the evidence. One question exposes the gap in a single line: on what examinable basis did you establish that this score means this concept?
Four further questions can be put before buying any instrument of this kind. What is your written definition of the concept? Against which independent measure was the score compared? Against which neighbouring concept was it compared, to show that it stands apart from it? On whom was all of that carried out? A supplier who answers all four with a general account of why the concept matters has answered none of them.
What counts as evidence of construct validity
- Convergence. The score agrees with another, independent measure of the same concept. When two instruments claim the same thing and give divergent results, at least one of them is wrong.
- Discrimination. The score stands apart from what is a different concept. An instrument titled “motivation” whose score rises with everyone who expresses themselves well is measuring something other than what is printed on its cover.
- Expected behaviour. The score moves where the concept says it should. A measure of occupational burnout would be expected to separate a team working at the peak of its season from a team outside it.
Evidence on this route accumulates; no single check settles it. That is why a construct is said to be supported more than it is proven.
The three are examined together, not one at a time. Convergence alone is easy to satisfy: two instruments written in a similar style can converge for that reason alone. The case becomes strong when convergence with one concept and separation from another appear in the same group of people, because the separation is what rules out the easy explanation.
In the sources we reviewed, we found no published evidence of construct validity for any particular instrument, and no figures settling how much construct validity an instrument has. We set no threshold at which convergence counts as sufficient, none at which separation counts as sufficient, and no number of checks that is enough. Anyone offered an instrument by a vendor has the technical documentation for that instrument to examine, not a general description of the concept of validity.
A worked example of construct validity: convergence that reassures, discrimination that exposes
Eight applicants sat an instrument titled “persistence”, together with two other instruments: an independent measure of the same concept, and a measure of fluency in written expression. The scores, given in that order, were:
- Applicant 1: 88, 84 and 90. Applicant 2: 81, 79 and 86. Applicant 3: 76, 72 and 80. Applicant 4: 70, 74 and 71.
- Applicant 5: 66, 61 and 68. Applicant 6: 60, 64 and 63. Applicant 7: 54, 51 and 57. Applicant 8: 47, 50 and 49.
Look first at the persistence instrument and the independent measure. Both sets of scores fall together from the highest to the lowest, yet the rank order matches in only four places out of eight. Applicants 3 and 4 swap places between the two instruments, as do applicants 5 and 6, so half of the eight change position. Even so, a reader can take this as reassuring convergence, because the eye falls on how close the scores are rather than on the change in positions.
Then look at the persistence instrument and the fluency measure. The rank order matches in eight places out of eight, and not one applicant moves from their position. The instrument follows written fluency more closely than it follows the independent measure of the concept whose name it carries.
Had the first check been the only one, the conclusion would have been that the instrument is supported. The second check reverses that conclusion. What the score measures is closer to skill with words than to persistence, and anyone who writes well scores high on this instrument whether they persist or not. That is what it means to say that discrimination is the check that exposes, and that convergence alone satisfies anyone who wanted to be satisfied.
The scores above are assumed in order to show the difference between the two checks. They are not attributed to any existing instrument and are not the result of a study.
Construct validity does not travel with a translated instrument
One weakness specific to this route arises when an instrument is moved from one context to another by translating its wording. The items in such instruments are not pieces of information to be carried across. They are situations that every reader is assumed to understand in the same way.
An item describing a situation that is familiar in one working environment and unusual in another does not measure the same thing in both, even when the translation is faithful word for word. The same applies to an item built on speaking plainly, written where plain speaking is welcome and then answered where the same remark would count as rudeness. What it measures there is the willingness to say the remark, not the trait it was written for.
What is needed in that case is not a review of the translation but a repeat of the checks: convergence and discrimination among the people to whom the instrument will actually be given. Whoever moves the instrument and assumes the evidence moved with it has carried the questions across and left behind what supports them.
Where the absence of construct validity shows its effect
A weak instrument on this route does not produce a number that is visibly wrong. It produces a plausible number, which is reported under the concept’s name, and judgements about people are then built on it.
One example is a low score labelled “poor discipline” when it is in fact the effect of question wording that some applicants understood differently. The word in the title passes into the judgement on the candidate without passing through any evidence.
In integrity testing the problem is sharper, because the concept there is far removed from observation and the consequence of the judgement for the candidate is serious.
The effect does not stop at the individual applicant. An organisation that used an instrument tracking fluency of expression, believing it tracked a trait relevant to the work, will have narrowed its intake to one kind of person without intending it or knowing it. The results will not reveal this, because the scores will look consistent and graded, and the people screened out appear in no report once they have been screened out.
What is not evidence of construct validity
Some things can be offered in place of evidence on this route without being evidence:
- The items of the instrument hang together. Coherence among the items shows that they point in one direction. It does not show that the direction is the concept printed on the cover. That is a question about the consistency of the score, and it is the subject of test reliability.
- The concept matches what the textbooks say about it. Agreement with the theoretical description is a condition of the definition, not evidence about the measurement.
- Applicants see the instrument as connected to the concept. That is an impression formed from the surface of the questions, and it can be the strongest feature of a weak instrument, because whoever wrote it knew which words suggest the concept.
Construct validity attaches to a reading of the score in a given use
The phrases “a valid instrument” and “an invalid instrument” compress the meaning until it is lost. What evidence supports is not the instrument but a particular reading of its score, in a particular group, for a particular purpose. Those are three conditions, and none of them can be dropped.
Suppose evidence supports reading an instrument’s score as persistence among new graduates, and the instrument is used to screen them at entry. That support does not extend to reading the score as persistence among people with ten years in work, nor to using it in a promotion decision. Both extensions can happen without anyone noticing, because the instrument is the same and the name is the same. What changed is who sat it and what was built on it.
That is why the fourth question in the list above, “On whom was all of that carried out?”, carries more weight than it appears to. The answer limits the support rather than widening it: it says where the established evidence applies, and by implication where it does not.
A further limit comes from another direction: the consistency of the score sets a ceiling on what the score can indicate. A score that swings for the same person between two sittings cannot indicate a stable trait in that person, because what fluctuates cannot indicate what holds steady. The relationship runs one way only. Consistency is necessary but not sufficient, and a perfectly stable score may be stable while measuring something else entirely. The worked example above showed this: the score that tracks written fluency tracks it steadily, and its steadiness does not make it persistence.
Reporting a score without overstating its construct validity
Part of this can be addressed in the wording of the report itself. Writing “the candidate is low in persistence” and writing “the candidate scored 47 out of 100 on instrument X” differ in more than politeness.
The first is a judgement about the person that attributes a fixed trait to them. The second is a fact about a measurement taken with a named instrument on a known date. Once written, the first travels between files and, two years later, looks like a description of the person. The second carries with it the questions that should be asked: which instrument? When? What is the evidence that this number means that trait?
Writing in the second form makes the claim examinable, and that is the whole of what is asked of the wording.
This is an explanation of the concept, not legal advice.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.