What overconfidence bias is
Overconfidence bias is a systematic gap between how accurate a judgement is and how much confidence is placed in it, with the confidence running higher. It does not say the judgement is wrong, only that the person holding it treats it as more accurate than it is, and so builds more on it than it can bear.
In HR it has two main homes: selection decisions, and estimates of figures for a coming period. It is more dangerous in the first because it goes undetected, and easier to measure in the second because the estimate can be set against what actually happened.
Three forms, and the third is the costly one
- Overestimating accuracy: believing one’s judgement is more correct than it is.
- Overestimating one’s standing: believing oneself better than others at something everyone does, such as judging people.
- Overprecision: giving far too narrow a range for a value one does not know. This is the least discussed and the most consequential at work, because it enters every figure estimated for a future period, and it is the only one of the three that can be measured without forming an opinion about anyone.
The calibration test: arithmetic, not opinion
Overprecision is measured with a simple procedure: for every estimate, ask for a range rather than a single number, together with a stated level of confidence in it, and then compare the range with what happened.
Take someone who estimates the workforce cost for each quarter, gives a range every quarter, and says they are 90% confident in it. That means they expect the actual value to fall outside the range about 10% of the time. Over 4 quarters, the expected number of misses is 0.4, so in most cases every quarter should land inside its range.
If the actual value falls outside the range in 3 quarters out of 4, the hit rate is one in four, or 25%, for ranges given at ninety per cent confidence. The fault is not in the estimate itself or in the circumstances. It is in the width of the range, which should have been much wider.
Two practical points follow. First, someone who gives a narrow range looks more expert than someone who gives a wide one, while the second is the more honest. Second, a decision built on a narrow range is taken with no plan for the figure falling outside it, because falling outside was never considered.
There is a caution in reading this test: a miss is not an error every time. A range stated at ninety per cent is supposed to miss occasionally, and someone whose actual figures never fall outside their range even once in twenty quarters is giving ranges wider than necessary. That is the opposite fault: an estimate that tells you nothing because it accommodates every possibility. Calibration is required in both directions, and the aim is for the hit rate to approach the stated confidence, not to reach one hundred per cent.
The figure is read over a series, not a single occasion. One quarter outside the range says nothing; three out of four says something to act on with care; twenty quarters say something to build on. This imposes one practical requirement: stated estimates must be kept somewhere they can be found again, because the whole test depends on what was said earlier being on record in writing.
The quarters, the ninety per cent and the three misses in this example are assumed purely to show how the test works. Four quarters is a small number on which no final judgement can rest; it is used only to show how what was stated is compared with what occurred, and the longer the series of quarters, the more reliable the judgement on calibration becomes.
How it shows in a hiring decision
The sharpest form of overconfidence in recruitment runs like this: an interviewer’s confidence rises with the amount they know about a candidate, and not everything they know bears on performance. A long, unstructured interview produces a rich and detailed impression, and that richness is itself what raises the confidence, even when most of it predicts nothing.
This is what separates it from the neighbouring patterns, and the difference is one of location, not degree:
- Rater biases such as leniency, severity and central tendency error make a rating reflect something in the rater or the instrument. They all concern where the rating lands. Overconfidence concerns how much weight the rating is given, and a rating can land exactly where it should while the reliance placed on it is still excessive.
- Implicit bias is a judgement that precedes observation because of membership of a group. Overconfidence comes after all the observation, as the weight given to it.
- The halo effect is an exaggerated generalisation from a real observation. It is the closest neighbour; the difference is that the halo spreads one trait across the other traits, while overconfidence spreads confidence across the whole judgement once it has formed.
The practical consequence is that the remedies for those patterns do not, on their own, address this one. Written criteria and a rating scale correct where the rating lands, and a performance calibration session aligns how raters use the scale, but neither says anything about how much should be built on the result.
Why it does not correct itself
The reason is that the feedback loop is broken, and the break is structural rather than a matter of neglect:
- Rejected candidates are never seen at work. The organisation does not know what the people it turned away would have done, so a wrong rejection never comes to light. The same problem appears on the measurement side of selection: an organisation sees at work only the people it accepted, which narrows the range of scores against which any fit can be measured.
- Accepted candidates’ performance is read as confirming the decision, when what shows may be the effect of employee onboarding or of the person who managed them, not of the judgement made at selection.
- The gap between decision and outcome is long. Months pass between the interview and anything known about the hire’s performance, and in that time the details of the original judgement are forgotten, leaving nothing to compare against.
That is why, in this respect, someone who has run two interviews and someone who has run a thousand are on the same footing, unless the thousand judgements were each compared with their outcomes. Experience alone raises confidence without raising accuracy when news of what happened never reaches the person judging.
What closes the loop
- Record the estimate before the outcome. Before the decision, the interviewer writes down an expected performance rating for the candidate and how confident they are in it, and that record is kept.
- Compare it later with quality of hire. That measure is where the loop closes; without it, selection judgements are never set against anything.
- Fix the procedure to reduce what enters the judgement without predicting anything, using the fixed questions of a structured interview and independent scoring before any discussion. Independence before discussion matters here in particular, because the first opinion voiced in the room raises the others’ confidence in it, not their accuracy.
- Ask for a range with every numerical estimate, and at the end of the year count how often the actual figure fell outside it. The procedure costs nothing, and its result, a single number, changes an estimator’s behaviour more than any warning does.
A rule that holds for rater bias applies here word for word: telling raters that bias exists does not remove it, and what reduces it is a change in the instrument and the procedure, not in intentions. All four steps above are of that kind. None of them asks anyone to be less confident; each asks that what was said be compared with what happened.
Other places it works in HR
- Estimating time to fill a vacancy. The last stage of the hiring process is the one the employer controls least, and it takes the largest share of the time. A promised start date is usually given without a range, as an estimate most of which lies outside the estimator’s control. Saying a role will be filled in six weeks gives a single number for a process with a stage the estimator does not control.
- Estimating the effect of an intervention before rolling it out. The figure from a small pilot is usually presented as a single number, not a range, and spending is then built on it. The range here is not caution; it is the honest information, because a pilot on a small group does not produce a single number in the first place.
- Judging a programme’s effect after it has run. This is the most hidden of the three, because the person who designed the programme usually evaluates it, so confidence in the judgement combines with an interest in the result. The remedy lies partly outside this bias: write down, before starting, what will count as success and what as failure, so the judgement becomes a comparison rather than an estimate.
We have not attached any magnitude or prevalence to overconfidence bias in organisations, and in the sources we reviewed we found none that estimates one.
What the label should not be stretched to cover
It does not mark someone as less competent. It occurs in the most capable people, and expertise in a field can even increase it where feedback is broken. Overconfidence is also different from decisiveness: a quick decision taken while acknowledging that the information is incomplete has nothing of overconfidence in it, since it is a choice made with its limits in view.
And it cannot be pinned on any single estimate. It describes a pattern that appears over a set of estimates, and no one estimate can convict anyone of it. Someone whose single estimate fell outside its range has had nothing established against them.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.