What the John Henry effect is
The John Henry effect is a change in the behaviour of the comparison group in a trial, caused by that group knowing it is the one being compared against. Its members put in more effort, the gap between the two groups shrinks or disappears, and the smaller gap is then read as a verdict on whatever was being tried.
The name comes from a widely told story about a labourer who raced a machine that was meant to replace him, beat it, and died from the effort. The story illustrates the idea and is not evidence for it. What matters is the meaning the term has settled into when trial results are interpreted.
It is the other half of the problem the Hawthorne effect describes. The Hawthorne effect concerns people who change because they know they are being measured; the John Henry effect concerns people who change because they know they are the benchmark. The first sits on the side where something was tried, the second on the side it is measured against.
Why it is harder to see than the Hawthorne effect
Attention in a trial goes to the side where something was tried, because that is where the question lies. The other side is treated as a fixed background: its number is taken and subtracted.
That assumption is exactly what breaks here. The comparison side is not inert material. It is people who know their performance will be shown next to someone else’s, and who may know that whatever is being tried will be rolled out to them if it succeeds. So they move, and nobody is watching them.
News also travels through an organisation by routes other than the official one. An organisation that never announced that one branch is being compared with another may find that the staff there knew in the first week, while the report is still written as though they did not.
A worked example: where the effect goes
Take an organisation that tried a new arrangement in one branch and measured it against another branch. Before the trial, each branch produced 100 units a day. The figures are assumed, chosen to show the method of calculation and the direction of the effect; they are not measured values.
- The branch where the arrangement was tried reached 112, a rise of 12.
- The comparison branch reached 108, a rise of 8, although nothing was applied there.
So the measured gap between the branches is 4 units. Had the comparison branch stayed at 100, as anyone reading the report assumes, the gap would have been 12 units.
The measured gap is therefore one third of what a fixed baseline would have shown. A decision built on the 4 units may be to drop the arrangement as not worth it, which means dropping an arrangement whose effect was 12 units.
This is the direction that sets the effect apart. The Hawthorne effect makes a result look larger than it is; the John Henry effect makes it look smaller. The first leads to rolling out what does not deserve it, the second to neglecting what does. The second is the less visible of the two, because nobody follows up the effect of something that was dropped.
The two effects do not cancel out
It is sometimes said that the two effects offset each other: one lifts one side, the other lifts the other, so the gap survives intact. That argument rests on a premise nobody has established, which is that the two are the same size.
Nothing requires that. The side where something was tried receives a new tool, training and attention from management; the comparison side receives nothing but the knowledge that it is being compared. The two causes differ in kind, and there is no reason for them to be equal in size. We found nothing, in the sources we reviewed, that establishes how the John Henry effect relates to the Hawthorne effect in size, which is one more reason not to assume they balance.
It follows that a measured gap can be larger than the true one, smaller than it, or equal to it by coincidence, and the gap itself does not say which. The right response is not to estimate a correction. It is to show separately the two figures by which each side moved, so the reader can see that the comparison side moved at all.
Where it arises in HR work
When something in HR is tried on one part of an organisation and measured against another, the arrangement itself leaves the door open to this effect:
- A new tool or procedure in one branch or department. The comparison department knows its result will be shown beside someone else’s. This is the most obvious setting.
- A new pay or incentive arrangement, such as a change to the salary structure in one unit. Here the effect is stronger, because the comparison side does more than know it is measured: it watches others receive something it does not. Its effort can move either way, up to prove it does not need the arrangement, or down out of a sense that what was handed out passed it by.
- A training programme offered to one team and not another. This is the most contaminated case, because the team left out may seek what it was denied by other means, getting trained outside the programme while being counted as untrained.
- A new performance measure, such as new performance standards for one group. Anyone who knows they will be compared with people judged on a different measure pushes up whatever they themselves are measured on, and the comparison becomes one between two arrangements that both moved.
What the four share is that the comparison side has a stake in the result: the outcome may be rolled out to it, or read as a judgement on its own performance. Someone with a stake in the result of a measurement cannot be measured as a fixed baseline.
How to detect it after the trial
We found no source, among those we reviewed, that sets a size for the John Henry effect, a period over which it lasts, or the conditions under which it appears. Four signs are weighed together, and none of them proves anything on its own:
- Read the comparison side’s line by itself. If it moved away from its previous level during the trial, something needs explaining. This is the first thing to check, and it is easy to skip when a report shows the gap rather than the two lines.
- Ask what was known. Did the staff there know, and when did they find out? If the knowledge arrived halfway through, it shows as a break in the line at that point.
- Look at what happened after the trial ended. If the comparison side drops back to its earlier level once the trial is over, that is strong evidence, because it is movement that only the removal of its cause explains.
- Read measures unrelated to the trial. A rise on the comparison side across several key performance indicators at once looks like general extra effort rather than improvement in any particular piece of work.
The label fits only when knowing about the comparison is the only plausible thing that changed. The comparison side can also rise because of the season, a shift in the market, or another decision taken in the organisation during the same period, and none of those is this effect.
What the report must record to stay readable
What is lost here can be lost in the writing rather than in the measuring. A report that shows only the gap prevents its reader from seeing what they need, even if every number is safely held by whoever wrote it. Three items belong in every such report:
- Each side’s value before and after the trial, shown separately, rather than the gap between them. The gap is a derived number; the two originals are what show that the comparison side moved.
- What was known and what was not. One sentence saying who was told what, and when. Without it the report cannot be reread a year later, because its author will have forgotten.
- The measurement period, and when it began on each side. If measurement on the comparison side started after its staff knew, its baseline is gone.
The first item can settle the argument, because it turns the question from whether the trial worked into why the side where nothing was tried moved, and that is a question with an answer to look for.
The analytical tools used on these figures describe what happened on both lines, but without information about who knew of the comparison they cannot separate the reason one line moved from the reason the other did.
Design choices that reduce it
- A baseline from the past. Compare the branch where the arrangement is tried with itself over a similar earlier period, instead of with a branch that knows. The limit is that whatever changes between the two periods, season or market, enters the gap.
- Rollout in waves. Extend the arrangement to sites one after another, so each site is its own baseline before its turn and nobody spends long as the comparison.
- Distance between sites. Choose a comparison side the news does not reach easily. This helps less the smaller the organisation, and not at all if both teams work in one building.
- Not labelling anyone the comparison group. The label itself produces the effect. It is enough to say that an arrangement is being tried in one place and that the results will be shared, without naming anyone as the other side of a comparison.
These describe practices that reduce the effect; they are not a settled methodological rule, and anyone who wants the detail should go to the literature on experimental design itself.
Hiding the trial from staff is not an open option. What is collected about them during it, and any decision built on it, concerns them, and it is personal data that the rules on employee data reach. The point is to control what is said, not to conceal that something is going on.
Nor is the effect an argument against running trials. The alternative to a trial affected by it is not a decision taken without a trial; it is a trial in which both lines are read together. The same caution applies to the Hawthorne effect: the aim is to control how a number is read, not to do without the number.
A third effect that gets mixed up with both
Something with a third cause is sometimes attributed to these two effects. A site whose performance was unusually low is picked for the trial, and it then climbs back towards its normal level with no help from what was tried, the pattern usually called regression to the mean. It differs from both effects in that nobody needs to know anything; it happens even if the trial is kept completely secret.
Its sign in the data is plain: the site’s line before the trial was erratic, and the starting point fell at the lowest point it reached. Protection against it comes not from design alone but from the choice of site. Do not pick the branch that performed worst in the previous month for the trial, which is precisely what someone hoping to show an effect tends to do.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.