Understanding OLS Regression: How Economists Estimate Real Relationships
- ▸OLS fits the line that minimizes squared errors; under the Gauss–Markov conditions it is the best linear unbiased estimator.
- ▸Read a regression through the slope, its standard error and t-statistic, and its economic size; R² alone says little about usefulness.
- ▸Omitted variable bias makes OLS precise but wrong when a left-out factor is correlated with the regressor, and more data does not fix it.
Understanding OLS regression — simply put
Ordinary least squares (OLS) estimates how much one variable changes, on average, when another changes by one unit. It fits the straight line through a cloud of data points that makes the squared vertical distances from the points to the line as small as possible. Almost every empirical claim in economics, from the return to a year of schooling to the effect of a rate rise on inflation, starts with an OLS regression.
Why square the errors rather than add them?
Positive and negative errors would cancel if simply added, so the line must penalize both. Squaring does that, gives a unique solution with a closed form, and penalizes large misses more than small ones. The payoff is the Gauss–Markov theorem: when the errors have mean zero, constant variance and no correlation with each other or with X, OLS is the best linear unbiased estimator. No other linear, unbiased method gives estimates with smaller variance.
How do you read a regression table?
The table regresses weekly earnings on years of schooling for ten workers. The sample is a teaching example built to resemble U.K. survey patterns; every statistic is computed exactly from it.
| Term | Coefficient | Std. error | t-statistic |
|---|---|---|---|
| Intercept (β₀) | -43.80 | 33.25 | -1.32 |
| Years of schooling (β₁) | 44.04 | 2.39 | 18.42 |
The slope says each extra year of schooling is associated with £44 more a week. Its standard error of 2.39 measures sampling uncertainty; the t-statistic of 18.4 is far above 2, so the slope is clearly different from zero. The intercept of −£43.80 is an extrapolation to zero years of schooling, outside the data, and has no meaning on its own. An R² of 0.977 says schooling accounts for almost all the variation in this small sample; real survey data, with thousands of noisy observations, typically give R² values far lower without the slope being any less useful.
When does OLS give the wrong answer with confidence?
The most dangerous failure is omitted variable bias. If a variable that affects Y is left out and is correlated with X, its effect is loaded onto the slope:
Ability is the classic example. More able people tend to stay in school longer and also earn more, so a regression of earnings on schooling alone overstates the return to schooling. A larger sample does not fix this: the estimate converges to the wrong number. Fixing it requires a different design, such as instrumental variables, panel data or a natural experiment.
What does a regression on economy-wide data look like?
Okun's law is one of the most famous OLS relationships in macroeconomics: when real GDP grows faster than its trend, unemployment falls. Regressing the change in the unemployment rate on real GDP growth for the United States has historically produced a slope of roughly −0.5: each extra point of growth goes with about half a point less unemployment.
Time-series regressions like this need extra care. Errors are usually correlated over time, which makes ordinary standard errors too small; economists use heteroskedasticity- and autocorrelation-robust (Newey–West) standard errors instead.
The takeaway
OLS gives the best linear unbiased estimate when its assumptions hold, and a precise, confident, wrong one when an omitted variable is correlated with the regressor. Read the slope, its standard error and its economic size, and always ask what the regression has left out.
Further reading
- Time Series Analysis Explained — internal
- Difference-in-Differences Explained — internal
- FRED: Unemployment rate (UNRATE) — external reference
Frequently Asked Questions
What does OLS regression do?
Ordinary least squares estimates the relationship between an outcome and one or more explanatory variables by choosing the line that minimizes the sum of squared differences between observed and predicted values. The slope measures the average change in the outcome for a one-unit change in the explanatory variable.
What is the Gauss–Markov theorem?
The Gauss–Markov theorem states that when regression errors have mean zero, constant variance and no correlation with each other or with the regressors, OLS has the smallest variance among all linear unbiased estimators. OLS is then called BLUE, the best linear unbiased estimator.
What is omitted variable bias?
Omitted variable bias arises when a variable that affects the outcome is left out of a regression and is correlated with an included regressor. The included coefficient then absorbs part of the omitted effect, and the bias does not shrink as the sample grows.
What is a good R-squared?
There is no universal threshold. R-squared measures the share of variation in the outcome explained by the regressors; cross-section data on individuals often give low values even when coefficients are precise and important, while trending time series can give high values from spurious relationships.
Primary Sources
Cite This Article
EconoLens Research Desk. (2026, October 11). Understanding OLS Regression: How Economists Estimate Real Relationships. EconoLens. https://www.econolens.co.in/news/study-ols-regression-explained
The EconoLens Research Desk reviews academic papers in economics and econometrics, translating cutting-edge research into accessible analysis. Full credit is given to original authors in every review.