Ennio Alessandrini

Chapter I

Introduction to Regression Analysis

01What Regression Analysis Is

Regression analysis is a statistical methodology used to study the relationship between a dependent variable and one or more independent variables.

The dependent variable is the outcome we want to explain or predict (usually denoted by Y). The independent variables are regressors/predictors (X in simple regression, X1, X2, ..., Xk in multiple regression).

Linear regression conceptual diagram

Central question: How does the expected value of Y change when one or more explanatory variables change?

  • Prediction: estimating future or unknown values of Y from observed X.
  • Explanation: quantifying direction and magnitude of association.
  • Inference: testing whether observed relationships are statistically meaningful.
  • Decision support: guiding economic, business, scientific, and policy choices.

Regression is a formal framework that decomposes observed variation in Y into a systematic component (explained by regressors) and a random component (noise, omitted factors, measurement error).

Chapter II

The Population Regression Model and Its Interpretation

02The True Population Regression Model

In simple linear regression, the theoretical population model is:

Eq. 1

Meaning of each term:

  • Yi: value of dependent variable for observation i.
  • Xi: value of independent variable for observation i.
  • beta0: intercept parameter.
  • beta1: slope parameter.
  • epsiloni: error term for observation i.

Interpretation of the intercept beta0:

Eq. 2

Interpretation of the slope beta1:

Eq. 3

The slope effect is constant in simple linear regression: the conditional expectation of Y changes linearly with X.

The error term collects omitted explanatory variables, measurement error, unobserved heterogeneity, random shocks, and model misspecification.

03Conditional Expectation Form of the Model

A formal representation of the regression model is:

Eq. 4

The observed value differs from this conditional mean by the error term:

Eq. 5

Substituting gives the population regression equation:

Eq. 6
  • Deterministic part: conditional expectation.
  • Stochastic part: deviation from conditional expectation.

04Why the True Model Is Unknown

In practice, the whole population is not observed. We observe only a sample:

Eq. 7

Therefore, population parameters beta0 and beta1 are unknown and must be estimated from sample information.

Estimated regression line:

Eq. 8
  • beta-hat-0 estimates beta0.
  • beta-hat-1 estimates beta1.
  • Y-hat-i is the predicted value for observation i.

Chapter III

Least Squares, Residuals, and the Fitted Line

05The Least Squares Principle

Ordinary Least Squares (OLS) chooses the fitted line that makes observed points as close as possible to the fitted values.

Residual definition:

Eq. 9
Eq. 10
Scatter plot with fitted regression line and vertical residual segments
  • ei > 0: point above fitted line.
  • ei < 0: point below fitted line.
  • ei = 0: point on fitted line.

OLS squares residuals to avoid cancellation of positive and negative deviations.

Eq. 11
Eq. 12
Eq. 13

OLS objective: choose beta-hat-0 and beta-hat-1 to minimize total squared residual distance.

Eq. 14

07The Fitted Line and Prediction

After estimating coefficients, the fitted equation is:

Eq. 28

Prediction for a new value X* is:

Eq. 29
  • The fitted line summarizes the relationship in-sample.
  • The same equation gives a prediction rule for new X values.

Chapter IV

Derivation of the OLS Estimators

06Derivation of the OLS Estimators

Objective function:

Eq. 15

First normal equation (differentiate with respect to intercept):

Eq. 16
Eq. 17
Eq. 18
Eq. 19
Eq. 20

Second normal equation (differentiate with respect to slope):

Eq. 21
Eq. 22
Eq. 23
Eq. 24
Eq. 25
Eq. 26
Eq. 27

Geometric property: the fitted least-squares line passes through (X-bar, Y-bar).

Chapter V

SST, SSR, SSE, and the Decomposition of Variation

08Decomposition of Variation: SST, SSR, and SSE

Total variation in Y is decomposed into explained variation and unexplained variation.

Total Sum of Squares (SST):

Eq. 30

Regression Sum of Squares (SSR):

Eq. 31

Error Sum of Squares (SSE):

Eq. 32

Notation used in these three formulas:

observed value of the dependent variable for observation i.

fitted or predicted value from the regression line for observation i.

sample mean of Y.

Conceptual variance decomposition visual showing SST, SSR, and SSE around the best fit line

Fundamental identity:

Eq. 33

Total variation = explained variation + unexplained variation.

09Why SST = SSR + SSE

Start from the identity:

Eq. 34

Square both sides:

Eq. 35

Sum over all observations:

Eq. 36

Under OLS, residuals are orthogonal to fitted values:

Eq. 37
Eq. 38
Eq. 39

Chapter VI

Measures of Fit: MSE, SER, R-squared, and Adjusted R-squared

10Mean Squared Error and Standard Error of Regression

Mean Squared Error (MSE) in simple regression:

Eq. 40

In multiple regression with k predictors:

Eq. 41

Standard Error of Regression (SER):

Eq. 42
Eq. 43
Eq. 44

SER measures the typical size of residuals in units of Y: small SER means points are close to the fitted line, large SER means they are more scattered.

Eq. 45
Eq. 46
Eq. 47
Eq. 48

11Coefficient of Determination: R^2

R-squared measures the proportion of total variation in Y explained by the regression model.

Eq. 49
Eq. 50
Eq. 51
Eq. 52
  • R^2 = 0: model explains none of variation.
  • R^2 = 1: model explains all variation.
  • 0 < R^2 < 1: partial explanation.
  • Example: R^2 = 0.60 means 60 percent explained and 40 percent unexplained.

Important caution: high R^2 does not automatically imply causality, correct specification, significant coefficients, or strong out-of-sample performance.

12Adjusted R^2

Ordinary R^2 never decreases when new regressors are added, even if they add little useful information.

Eq. 53
Eq. 54
  • n is sample size.
  • k is number of predictors.
  • Adjusted R^2 penalizes unnecessary regressors.
  • It can increase with useful variables and decrease with weak ones.
Eq. 55
Eq. 56

If SSE = 0, adjusted R^2 is also 1 because there is no residual error to penalize.

Chapter VII

Statistical Inference: Standard Errors, t-Tests, and F-Tests

13t-Test for an Individual Coefficient

Beyond estimation, regression supports inference: is the relationship significant, is slope different from zero, and is explanatory power beyond random noise?

For coefficient beta_j, test:

Eq. 57
Eq. 58
Eq. 59

General t-statistic:

Eq. 60
Eq. 61
Student t-distribution table diagram for hypothesis testing

Decision rule based on p-value with residual degrees of freedom (n-2 in simple, n-k-1 in multiple):

Eq. 62
Eq. 63

Large absolute t suggests stronger evidence against H0; small absolute t suggests weaker evidence.

Reference tableStudent's t Critical Values · Two-Tailed

Select the row corresponding to the appropriate degrees of freedom and the column corresponding to the chosen significance level .
Reject if .

dfalpha = 0.10alpha = 0.05alpha = 0.01
16.31412.70663.657
22.9204.3039.925
32.3533.1825.841
42.1322.7764.604
52.0152.5714.032
61.9432.4473.707
71.8952.3653.499
81.8602.3063.355
91.8332.2623.250
101.8122.2283.169
121.7822.1793.055
151.7532.1312.947
201.7252.0862.845
251.7082.0602.787
301.6972.0422.750
401.6842.0212.704
601.6712.0002.660
1201.6581.9802.617
inf1.6451.9602.576

14Standard Error of a Coefficient and F-Statistic

In simple regression, the slope standard error is:

Eq. 64
  • Higher residual noise (SER^2) implies less precision.
  • Greater variation in X implies more precision.

Overall model significance in multiple regression:

Eq. 65
Eq. 66
Eq. 67
Eq. 68
Eq. 69
Eq. 70
Eq. 71
Eq. 72
Eq. 73

In simple regression (one regressor), the overall F-test is equivalent to the t-test on the slope.

Chapter VIII

Multiple Regression and Overfitting

15Multiple Regression

Multiple regression model:

Eq. 74

Fitted multiple-regression equation:

Eq. 75

Each coefficient is interpreted holding all other regressors constant (partial effect).

Example interpretation: beta2 is the expected change in Y for a one-unit increase in X2, keeping other explanatory variables fixed.

Why useful: many outcomes are influenced by several factors simultaneously.

  • price
  • advertising
  • consumer income
  • seasonal effects
  • competitor actions

16Overfitting

Overfitting occurs when too many regressors are included relative to the informative structure in data; the model starts capturing random fluctuations instead of systematic relationships.

Overfitting regression visual comparing fit complexity

Consequences of overfitting:

  • very high in-sample R^2
  • excellent fit on historical data
  • poor performance on new observations

A strong in-sample fit does not guarantee strong out-of-sample prediction.

Adjusted R^2 helps because it penalizes regressors that do not sufficiently improve explanatory power; unlike ordinary R^2, adjusted R^2 can decrease when complexity is unjustified.