What the Hawthorne effect is
The Hawthorne effect is a change in how people behave because they know they are being measured, not because anything in their work has changed. The act of measuring enters what is being measured, and part of the result becomes an effect of being observed rather than an effect of the change the measurement was meant to estimate.
The name comes from a series of studies run at a factory known as the Hawthorne Works in the 1920s and 1930s. How the results of those original studies should be interpreted is still disputed, and we take no side in that dispute. The sense intended here is the one the term settled into in management use, not a view on what those studies showed.
One caution belongs at the start. The effect is not an argument for giving up measurement. Measuring without attention to it produces a number on which a wrong decision gets built; not measuring produces a decision with no number behind it at all. The aim is to control how the number is read, not to do without it.
The trial that worked and then did not repeat
Take an organisation that tried out a new tool for organising work with a single team of twenty employees. Output per person per day was 84 units before the trial and reached 96 units during the trial weeks: an increase of 12 units, or 14.3 percent.
The tool was then rolled out to 240 employees, and output settled at 87 units, an increase of 3 units, or 3.6 percent. The effect the business case rested on was exactly four times the real one: twelve units against three.
The budget shows this more plainly than the percentage does:
- The increase expected on the trial’s figure: 240 employees times 12 units, or 2,880 units a day.
- The increase that actually arrived: 240 times 3, or 720 units a day.
What was delivered is 25 percent of what was promised, and neither the tool nor the way it was used had changed. What had changed is that the twenty knew they were being measured, and the 240 knew nothing of the kind.
The figures in this example are assumed, chosen to show the shape of the problem. We do not set a percentage by which the result of any trial should be discounted for the Hawthorne effect, and no figure above should be read as one.
A snapshot against a curve
The effect does not last, and that is how it is found. Read the trial team’s output across the whole period rather than at its start:
- Week 1: 96 units, 14.3 percent above the starting level.
- Week 4: 92 units, 9.5 percent.
- Week 8: 88 units, 4.8 percent.
- Week 12: 87 units, 3.6 percent.
The settled value is 87, the same figure that appeared after the rollout. The trial was not wrong; it was read in the wrong week. The one practical rule to draw from this is that a trial’s effect is measured at its end, not at its beginning, and the period is extended until the number stops moving, not until it reaches the figure that was hoped for.
Not every gap between a trial and its rollout is the Hawthorne effect
This qualification has to be stated, because the term gets used as a ready explanation for every trial whose result failed to carry over. Four other causes produce the same gap:
- How the trial team was chosen. Trials are usually run on the most capable and most willing team, so part of the gap is a difference in the team rather than in the tool.
- Support that does not scale. One trainer for twenty people is not one trainer for 240, and the difference in outcome is a difference in support.
- Management attention. A trial is followed weekly from a more senior level. That alone moves the numbers, and it stops when the rollout begins.
- Novelty. A rise that comes with anything new and then fades. It sits close to the Hawthorne effect and differs in its cause: novelty is a response to the new thing itself, the Hawthorne effect a response to knowing one is measured.
One test separates them, and it runs after the rollout rather than before it: keep following the trial team itself. If that team also falls to 87 along with the rest of the organisation, the gap lay in the measurement and the attention. If it stays at 96 while everyone else reaches 87, the gap lay in the team or in its support, and the tool works where it is supported.
The same arithmetic exposes the selection effect. Suppose the organisation’s average before the trial was 78 units and the trial team stood at 84. The team was already 7.7 percent above average before anything was tried on it. Anyone who compares 96 with the organisation’s average of 78 arrives at 23.1 percent, most of which has nothing to do with the tool. The right comparison is the same team before and after, not the average of people who were never in the trial.
The effect also runs the other way
Just as the people in a trial raise their performance, so can the people who know they are the group the trial is being compared against. A control group whose members learn that they are the reference may increase their effort, the distance between the two groups narrows, and the tool’s effect looks smaller than it is.
So the error runs in both directions: inflation on the side of those measured in the trial, and deflation on the side of those they are compared with. A manager who guards only against the first may drop a useful tool because the difference looked slight. That is the costlier of the two mistakes, because it is never discovered: nobody learns what a cancelled project would have produced.
Where it appears in HR measurement itself
The effect is not confined to production trials. In HR measurement it turns up most often in three places:
- Attendance reviews. Announcing that attendance records will be reviewed raises punctuality for as long as the review lasts, so the review ends up describing how people behave while under review, not how they usually behave.
- Surveys. A survey run after it has been announced that managers will see the results is answered differently from one where nobody knows who reads the answers.
- Appraisal pilots. A department piloting a new appraisal arrangement documents its ratings with a care that does not outlast the pilot, and a quality that came from the care is credited to the arrangement.
That is why the Hawthorne effect is best understood as one case of a methodological error that recurs across HR analytics: a number that looks like evidence can be an effect of the way it was gathered.
Monitoring attendance and performance also has a regulatory side, and the term does not settle it. We make no statement about what the regulations require or permit when performance or attendance is monitored; that is a question for the provisions themselves.
Where it sits among its neighbours
The axis that sorts the neighbouring ideas is who changes: the person measuring, the person measured, or nobody at all.
Rater errors change the person measuring. Central tendency error and contrast error in appraisal are faults in the appraiser’s judgement, and they are treated by training appraisers and by giving each level of the scale a description that separates it from the next. The Hawthorne effect changes the person measured, so none of that treatment reaches it. The appraiser is reading correctly; what is being read has moved.
In the case of a false improvement, nobody changes and the number is simply wrong. The Hawthorne effect is the opposite case. The rise during the trial weeks really happened, and the units were really produced. The mistake lies in crediting them to the tool and in assuming they would last, not in their having occurred.
The same axis explains why the effect is no reason to stop running trials. A trial on one team still costs less than a rollout to 240 people, and the thing that needs correcting is how its result is read, which leaves the trial itself as the cheaper route to an answer.
The effect follows what people believe, not what they are told
A finer point: the strength of the effect does not follow the announcement alone. It follows what the person being measured believes will be done with the number. A measurement understood as a way to improve the tool changes behaviour a little; one understood as a way to decide who stays in the unit changes it a great deal.
Stating the purpose of a measurement is therefore more than a matter of transparency. It is a control on the instrument itself. Telling people what will be done with the number reduces the mistaken guess their behaviour would otherwise be built on. Measuring without stating the purpose produces the strongest form of the effect, because each person assumes the worst case and acts on it.
The same mechanism explains why a number differs between a measurement announced as covering the whole organisation and one announced as being about a single unit. People in the unit read the second as a verdict on their manager or on themselves, which is worth keeping in mind whenever engagement is measured team by team.
What reduces it
- A baseline from records that already exist. Taken from data collected routinely before the trial was announced, it gives a reference the announcement could not touch.
- A longer period. The effect fades with time, and measuring once the number has settled comes close to what the rollout will deliver.
- Measurement as routine rather than event. What is always measured shows little of the effect; what is measured once a year after advance notice shows it at its strongest.
- A second trial on another team. Repeating the trial on a team that was not chosen for its ability separates the effect of the tool from the effect of the team, at less cost than a rollout.
- Measuring what intention alone cannot move. An indicator that a deliberate push can shift within a single week is more exposed to the effect than one that moves only with a change in how the work is structured, which applies to key performance indicators chosen for a trial as much as to anything else. The first kind is read alongside the performance standards for the role, not on its own.
None of these amounts to a design for experiments. Experimental design is a field with its own specialists and its own tools, and what is set out here describes a source of error in reading a result, nothing more.
Before a trial’s result is rolled out
The most useful test before a rollout is to ask whether the trial team knew it was in a trial. If the answer is yes, the number may carry some of this effect, by an amount nobody knows, and the decision is put forward on that basis rather than as though the number were final.
Then two numbers are asked for before anything is approved: the value of the indicator at the end of the trial rather than at its start, and the difference between the end and the start. If that difference is large and falling, the effect is visible in it, and the decision rests on the settled value. If the period was too short for any fall to show, the answer is that the trial has not finished yet, not that its result is high.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.