What regression analysis is
Regression analysis is a statistical method for estimating the value of one numerical variable from another variable or from several, by fitting to the data an equation that assigns each explanatory variable a coefficient. The coefficient is the estimated change in the variable being explained when the explanatory variable rises by one unit, holding the other variables in the equation constant.
In HR it is the tool for moving from describing what happened to estimating a relationship between two things. That move is a step up from descriptive analytics, and it has a precondition: analysis cannot repair flawed data. Everything below starts after that precondition is met, and concerns what the equation says and what it does not.
A worked example with one variable
Take a department in which the relationship between length of service in years and monthly pay in riyals has been estimated, giving this equation. The figures are assumed to show how it is read:
Estimated pay = SAR 6,200 + SAR 310 × years of service
- At 3 years: 6,200 plus 930, or SAR 7,130.
- At 8 years: 6,200 plus 2,480, or SAR 8,680.
Three things can be read from this. First, 310 is the estimated difference in pay between two employees who differ by one year of service in this data. Second, 6,200 is the value of the equation at zero years, which may or may not mean anything here, because zero often lies outside the range of the data the equation was built on. Third, and most important, both numbers describe this group over this period. They do not carry over to another department, because the equation was not built on it.
R squared, and three wrong ways to read it
The equation usually comes with a number called the coefficient of determination, written R squared. Suppose it is 0.34 in our example. Its precise meaning: 34% of the variation in pay among the employees of this group is explained by the equation as fitted.
Three readings are common, and all three are wrong:
- That 34% of pay is caused by length of service. The number is not a share of pay or of any amount. It is a share of how much pay varies between people. A group whose pay is tightly bunched can produce a low value even when the relationship is strong.
- That the equation is right in 34% of cases. An equation is not right or wrong case by case. It gives an estimated value, not a verdict that holds or fails.
- That a high value proves the model is sound. A high value can be obtained by adding many variables to an equation fitted on few rows. It is a better fit to the existing data, not a better ability to estimate a new case.
This is the same mistake that affects many single HR figures: a number is read as whatever it resembles rather than as the division it actually is.
Against these three stands a mistake from the opposite direction: rejecting the result because the number is low. An equation with an R squared of 0.09 can carry a coefficient of real practical significance, because explaining little of the total variation does not rule out the variable having a substantial effect in magnitude. The two numbers answer two questions: the coefficient is about the size of the effect, and R squared is about how much of the spread the equation explains. Rejecting a large coefficient because R squared is small answers one question with the other.
What adding a second variable does
Add job grade to the earlier equation, and the service coefficient falls from 310 to, say, 95. That does not mean one of the two equations is wrong. They answer different questions:
- Without grade: how much does the pay of someone with one more year of service differ from others, whatever their grade? The answer is 310, and it includes the fact that people with longer service tend to be in higher grades.
- With grade: how much does the pay of two people in the same grade differ when one has a year more service? The answer is 95.
Holding a variable constant here is an operation on the equation, not on the organisation. The equation calculates as if grade were the same for everyone, while in reality grade and service move together. “Controlling for grade” therefore describes what the calculation did, not a state of affairs that exists.
This has a practical consequence for one of the tool’s most familiar uses, pay equity analysis: adding a variable that is itself affected by what is being examined hides the effect without removing it. If grades themselves are awarded unevenly, and grade is then held constant in the equation, the gap that remains is not the whole gap. It is only the part that did not pass through grade. The result is arithmetically correct and misleading if read as the complete answer.
Does a regression coefficient prove that one variable causes another?
A correlation between two indicators does not mean one causes the other; the familiar example is higher turnover in a department that works longer hours. The same holds for regression analysis. A coefficient is no stronger than a correlation on this point, even though it looks more precise because it comes as a number with a unit.
A coefficient is a relationship conditional on what was put into the equation, and everything left out keeps its effect. If some variable affects both sides and was not measured, its effect shows up in the coefficient of something else. That is why a causal reading of a coefficient assumes that everything that matters has been put into the equation, an assumption the equation can neither prove nor disprove.
Four uses in HR, and what each one demands
- The trend line in a pay structure. The relationship between pay and job grade is estimated, giving a line against which each individual salary is compared. It is the same logic as a compa-ratio, except that a compa-ratio measures pay against the midpoint of a range set in advance, while the equation derives its line from the data itself. If the salary structure is applied consistently, the two lines come out close, and the distance between them is itself information about how the structure is applied rather than about pay.
- Examining pay differences between groups. This is the pay equity question above, read together with the point that a pay differential may have a legitimate cause in the work itself. The equation separates what is explained by the variables included from what is not, and passes no judgement on the remainder.
- Estimating the effect of one variable on an indicator, such as the effect of time to fill on one of the organisation’s key performance indicators. This demands the most caution about causation, because the question being asked is itself a causal one.
- Estimating before the event, such as forecasting expected vacancies. An indicator like the vacancy rate is read in its context, and the equation does no more than put a number on that context; judging the number still depends on everything said above.
What these four uses have in common is that the equation narrows the question without closing it. It turns “why is our pay like this” into “what remains once grade and service are taken into account”, and the second is a question that can be investigated, while the first never was.
Before a result goes to management
- Enough rows for the number of variables. An equation with five explanatory variables on thirty rows fits itself to the particulars of those thirty, producing a high number that does not hold on other data. The common rule of thumb is to require a certain number of rows per variable. That is an accepted practice, not a fixed limit, and in the sources we reviewed we found none that establishes a single value for it.
- No estimating outside the range of the data. An equation built on service between one and ten years says nothing about 25 years, even if the software accepts the number and returns a result.
- A look at the outliers. A single row far from the rest can reverse a coefficient in a small group. In pay data this happens often, because of one role that resembles nothing around it.
A further limit has nothing to do with the arithmetic. Employee data is personal data, and the discipline that goes with it applies: a defined purpose, no more data than that purpose needs, and protection of what is held, with the practical rule that analysis runs at the level of the group, not the individual. That last rule matters more for this tool than for most, because the equation will accept a single employee’s row and return a number for them. The employee file is where that data lives, and how it is processed is read from the official sources that govern it; nothing about it is borrowed from the method described here.
The equation, its coefficients and every value in the examples are assumed to show how a result is read. They are not a model to adopt, and no coefficient given here is taken from any source.
Telling it apart from related tools
Comparing two averages leaves out everything else. Comparing two departments on average pay says nothing about what else they differ on, which is precisely what regression tries to take into account.
A regression equation is not, in itself, a forecast. The equation describes an existing group, and whether it is fit to estimate a case that was not in it is a claim that needs evidence of its own, the same principle behind test validity, where validity is a property of the inference and not of the instrument.
A descriptive indicator, by contrast, is not replaced by regression at all; it comes first. An equation presented before anyone has looked at the distribution hides what the descriptive view would have shown, including a group that is in fact two groups.
A standalone Saudi HR system
One employee file holding the contract, the documents and their expiry dates, the attendance record, leave, salary and end-of-service entitlements. End-of-service, overtime and leave-balance calculations are built into the system.
A standalone system on its own subscription. The connection to Qoyod Accounting is now available.